On this path · Foundations 5. Reading the Corpus Yourself
Reading the Corpus Yourself
Settle any usage question with corpus tools instead of guessing.
Learning outcomes
The last page said your intuition is a record of your exposure. The problem is that your exposure has gaps, and inside those gaps your intuition is simply silent or, worse, confidently wrong because it is echoing your first language. This page gives you a way to check. It teaches you to consult the collective exposure of millions of native speakers, recorded and searchable, so you can settle a usage question with evidence instead of a guess.
After studying this page, you can:
- Explain what a corpus is and why a phrase’s frequency is evidence about whether it is natural.
- Run the four core moves: the frequency duel, the collocate search, the register filter, and reading the concordance.
- Use these to settle a real “is it X or Y” question in under a minute.
- Recognize the cases where frequency is the wrong test, so you do not over-trust the count.
- Tell the difference between consulting a balanced corpus and naively counting web hits.
Before we dive in
Picture the argument you cannot win from inside your own head. You are about to write “make a decision,” but a doubt flickers: is it “make a decision” or “take a decision”? You have heard both. Your intuition, which is just your exposure talking, gives a muddy answer, because your exposure happened to include both and never tallied them. You could ask a native colleague, but they will often shrug and say “both sound fine,” because their intuition, while strong, is not a counter either.
There is a way out of this that does not depend on anyone’s memory. Somewhere there is a record of how tens of thousands of native speakers actually chose, in real sentences, written and spoken, across newspapers and conversations and academic papers. That record is a corpus, and learning to query it turns “I think” into “I checked.” For an advanced learner, this is close to a superpower: it makes you independent. You no longer need a native nearby to tell you what sounds right, because you can look at what natives, in bulk, actually did.
What a corpus is, and why frequency counts as evidence
A corpus is a large, carefully assembled, searchable collection of real texts. The two you will use most are the Corpus of Contemporary American English, known as COCA, which holds about a billion words of American English balanced across spoken, fiction, magazine, newspaper, and academic sources, and the British National Corpus, known as the BNC, the equivalent reference for British English. There is also the Google Books Ngram Viewer, which charts how often a phrase appears in printed books over time. Because this domain anchors on American usage, COCA is your default, with the BNC as the British comparison.
The reason a corpus is useful rests on one idea: in a balanced corpus, frequency is evidence of convention. Naturalness, from the very first page of this domain, is “the version a native would choose.” A corpus lets you see that choice made at scale. If thousands of writers reached for “make a decision” and almost none wrote “do a decision,” that lopsided count is the conventional choice, made visible. You are not asking whether a phrase is grammatical, you are asking whether it is what the community actually says, and frequency answers exactly that.
Four moves you can make with a corpus
You do not need to be a linguist to get value from a corpus. Four simple moves cover almost every question an advanced learner has. The table names them, the question each answers, and a one-line example.
The four corpus moves, the question each one answers, and an example query.
| Move | Answers the question | Example |
|---|---|---|
| Frequency duel | Of two phrasings, which is conventional? | make a decision vs do a decision |
| Collocate search | Which words habitually go with this one? | what adjectives precede concern? |
| Register filter | Where does this phrase live, speech or academic prose? | is kick off too informal for a report? |
| Concordance read | What is the real pattern around this word? | 20 real sentences using comprise |
The four subsections take each move in turn. None takes more than a minute once you know the site.
The frequency duel
The frequency duel is the move you will use most. You have two candidate phrasings; you search each and compare the counts. The phrasing with the dominant count is the conventional one. This is how you settle “make versus do a decision” in seconds: one returns thousands of hits, the other essentially none.
The chart below shows the shape of a typical duel, for which verb goes with “decision” in American English. The heights are illustrative of the pattern a corpus search reveals, not exact counts, but the relationship is real and stable.
The middle bar is the lesson within the lesson. “Take a decision” is not an error; it is the normal British choice, much rarer in American English. A frequency duel only means something against a corpus of the variety you are aiming at, which is why you pick COCA for American English and the BNC for British, rather than mixing them. We return to this split throughout, beginning in collocation, the heart of word choice.
The collocate search
The second move asks not “X or Y” but “what goes with X.” Corpus tools can list a word’s collocates, the words that appear next to it far more often than chance, ranked by how strongly they bond. Search the collocates of “concern” and you find “raise,” “address,” “valid,” “growing,” “serious.” That list is a ready-made set of chunks to collect, straight from the previous page’s method. The collocate search is, in effect, an automatic chunk-finder for any word you care about.
The register filter
The third move answers “where does this phrase belong.” Because COCA tags every text by genre, you can see whether a phrase clusters in speech, in journalism, or in academic writing. This settles register questions directly. Wondering if “kick off” is too informal for a written report? Filter by genre: if it lives overwhelmingly in spoken and informal sources and barely appears in academic prose, you have your answer. This turns the abstract register dials of the previous track into something you can measure.
Reading the concordance
The fourth move is the richest and the most overlooked. A concordance is a list of every occurrence of a word, each shown in a line of its real context. Reading twenty concordance lines for a word teaches you its pattern faster than any dictionary definition, because you see the company it keeps, the structures it sits in, and the register it favors, all at once. If you are unsure how “comprise” is actually used, do not read its definition; read forty real sentences containing it, and the pattern (“the whole comprises the parts,” not “is comprised of”) becomes visible.
When frequency is not the answer
A corpus is powerful, but treating frequency as a verdict on everything will mislead you. Three boundaries keep it honest.
First, frequency is not the same as correctness. Common errors are, by definition, frequent; “could of” appears thousands of times and is still wrong. Frequency tells you what is conventional among the corpus’s writers, which usually but not always tracks what is correct. For genuinely disputed points, a usage guide that weighs the evidence, rather than a raw count, is the better authority, which is why this domain also leans on references like Garner.
Second, raw web counts are not corpus counts. Typing a phrase into a search engine and reading the hit number is tempting and badly flawed: the web is full of non-native writing, the counts are unreliable estimates, and there is no register control. A real corpus is balanced and clean; the open web is neither. Use the corpus for evidence and the web only for a rough sanity check.
Third, corpora lag. New expressions and very recent shifts may be underrepresented in a corpus assembled years ago, so a low count for a clearly current phrase may reflect the corpus’s age, not the phrase’s status. For fast-moving informal language, recency matters, and the corpus is a slightly old photograph.
“I googled it and got two million results” is not corpus evidence. Search-engine counts are unreliable, uncontrolled for register, and full of non-native text. When you want real evidence, use a balanced corpus such as COCA; keep web searches for a quick gut check only.
Mental Model: a second opinion, not an oracle
The wrong way to hold this tool is as an oracle that pronounces every phrase right or wrong. That leads to two failures: trusting the count even when it is measuring the wrong thing, and freezing every time you write because you feel you must verify everything.
The better model is the corpus as a second opinion, the kind you get from an expert colleague. You consult it when your own judgment is genuinely uncertain or when the stakes are high enough to be worth a check. You weigh what it says against what you know, especially about register and recency. And then you decide. Most of your English should still flow from intuition; the corpus is for the handful of moments a day when intuition shrugs and you want evidence. Used that way, it accelerates the noticing loop from the previous page rather than replacing your judgment with a number.
You are unsure whether to write 'different from' or 'different than' in a formal report for an American audience. What is the best single move?
Common mistakes
The first mistake is counting web hits and calling it evidence. Search-engine numbers are noisy, uncontrolled, and polluted by non-native text. Use a real, balanced corpus for any claim you intend to rely on.
The second mistake is ignoring the variety. A frequency duel run against a British corpus answers a British question. If you are aiming at American English, use COCA, and read a high British-only count, such as “take a decision,” as a British signal, not an error.
The third mistake is ignoring register. A phrase can be very frequent overall and still be wrong for your situation because it lives in the wrong genre. Always ask not just how often, but where, by filtering the corpus by genre.
The fourth mistake is verifying everything. The corpus is for genuine uncertainty and high stakes, not for every sentence. Over-checking freezes your writing and trains dependence instead of intuition. Let most of your English flow, and reach for the corpus only when you would otherwise be guessing. With the foundations now in place, the rest of the domain applies them, starting with the grammar that still trips up advanced speakers and the vocabulary that carries the most naturalness, in collocation, the heart of word choice.
Mastery Questions
Question
What is a corpus, and why does a phrase's frequency in one count as evidence about naturalness?
Answer
A corpus is a large, carefully balanced, searchable collection of real texts, such as COCA for American English (about a billion words across spoken, fiction, magazine, newspaper, and academic sources) and the BNC for British English. Frequency counts as evidence because naturalness is the version a native would choose, and a balanced corpus shows that choice made at scale: if thousands of writers used one phrasing and almost none used the alternative, the lopsided count is the conventional choice made visible. It measures convention, not grammaticality.
Sources & evidence5 claims · 6 cited
Corpus descriptions and the frequency-as-evidence principle are grounded in COCA, the BNC, Sinclair, and Biber; the make/take/do decision pattern is grounded in COCA, Oxford Collocations, and Garner; the frequency-chart heights are explicitly illustrative, not measured counts.
- COCA (the Corpus of Contemporary American English) is a balanced reference corpus of about a billion words spanning spoken, fiction, magazine, newspaper, and academic registers, and the BNC is the comparable British reference corpus.verified
- In a balanced corpus, the relative frequency of a phrasing is evidence of its conventionality, because it reveals the choice native writers and speakers actually made at scale.verified
- Make a decision is the dominant collocation in American English, take a decision is common in British English, and do a decision is essentially unattested in either variety.verified
- COCA tags texts by genre, which lets a user filter a query by register (spoken, journalistic, academic) and read concordance lines showing each occurrence in real context.verified
- Frequency is not identical to correctness: common errors are frequent, raw search-engine counts are noisy and uncontrolled for register, and corpora assembled years earlier underrepresent very recent usage.verified
Cited sources
- Corpus of Contemporary American English (COCA) · Mark Davies, english-corpora.org
- British National Corpus (BNC) · english-corpora.org
- Corpus, Concordance, Collocation · John Sinclair
- Longman Grammar of Spoken and Written English · Biber, Johansson, Leech, Conrad, Finegan
- Oxford Collocations Dictionary for Students of English, 2nd edition · Oxford University Press
- Garner's Modern English Usage, 4th edition · Bryan A. Garner