CACrown ArchivesThe cinema collection
Menu
Research dossier · General Reference

Corpus of Contemporary American English

a more than 560-million-word corpus of American English

Cross-disciplinary reference desk with index cards, atlas, dictionary and catalogue
General referenceInterpretive dossier study · Crown Archives visual atlas
Record originEnglish Wikipedia
Text licenseCC BY-SA 4.0
Source revisionJun 2, 2026
Entity authorityQ5172683 ↗
Source-derived summary

The Corpus of Contemporary American English (COCA) is a one-billion-word corpus of contemporary American English. It was created by Mark Davies, retired professor of corpus linguistics at Brigham Young University (BYU).

Content

The Corpus of Contemporary American English (COCA) is composed of one billion words as of November 2021. The corpus is constantly growing: In 2009 it contained more than 385 million words; in 2010 the corpus grew in size to 400 million words; by March 2019, the corpus had grown to 560 million words.

As of November 2021, the Corpus of Contemporary American English is composed of 485,202 texts. According to the corpus website, the current corpus (November 2021) is composed of texts that include 24-25 million words for each year 1990–2019.

For each year contained in the corpus (1990–2019), the corpus is evenly divided between six registers/genres: TV/movies, spoken, fiction, magazine, newspaper, and academic (see Texts and Registers page of the COCA website). In addition to the six registers that were previously listed, COCA (as of November 2021) also contains 125,496,215 words from blogs, and 129,899,426 from websites, making it a corpus that is truly composed of contemporary English (see Texts and Register page of COCA).

The texts come from a variety of sources:

TV/Movies subtitles: (128 million words) Texts taken from the OpenSubtitles collection of American TV shows and movies.

Spoken: (127 million words) Transcripts of unscripted conversation from nearly 150 TV and radio programs.

Editorial summary

This brief starts where responsible research should: with the source description of “Corpus of Contemporary American English” as a more than 560-million-word corpus of American English. Everything that follows is an evidence route, not borrowed authority.

Editorial reviewA dependable orientation record for establishing vocabulary, names and a first evidence trail. The current lead gives the account dated anchors—2021, 2009, 2010, 2019—that can be checked directly. The selected authority fields contribute no independent date. The account is most persuasive where Corpus, Contemporary and American can be independently traced.
Editorial analysis

Why this record matters

The subject matters to the general reference register because the source frames it as a more than 560-million-word corpus of American English. Its deeper value depends on whether names, dates, institutions and citations support that framing.

Evidence profile

The citation trail is more important than the brevity of the summary: it shows where individual claims can be examined in context. The source revision retrieved here is dated Jun 2, 2026. The linked authority identifier is Q5172683. None of the 1 selected statements returned an explicit reference. The first chronological checks are 2021, 2009, 2010 and 2019.

Critical limits

Overview language is designed for orientation and should not be treated as a substitute for the evidence cited beneath it. The lead is largely declarative, so disagreement and counter-evidence require a deliberate search beyond the opening account. Authority statements aid reconciliation but still require their own references, qualifiers and ranks to be checked.

How to read it

Use the entry as an orientation point, then follow its citations and revision history. Names, dates and institutional relationships should be checked against the original record.

Best used for
  • Subject orientation
  • Search vocabulary
  • Locating named sources
Verify next

The closest primary source, responsible institution and strongest cited specialist reference.

Three-step research path

  1. Establish the record: confirm the title “Corpus of Contemporary American English”, its source revision and the description used here.
  2. Expand the search: follow Corpus of Contemporary American English primary sources, Corpus of Contemporary American English archive and Corpus research across catalogues and specialist indexes.
  3. Test the account: compare the strongest cited source with the responsible institution’s current record and note any disagreement.

Questions for further research

  1. Which source most directly establishes the central claim about “Corpus of Contemporary American English”?
  2. Which institution is responsible for the underlying evidence?
  3. Which cited source is closest to the event, object or claim?
Subject index

Search terms from this dossier

Source & attribution

This entry incorporates text from “Corpus of Contemporary American English” on English Wikipedia. Contributors are listed in the page history. Text is available under the Creative Commons Attribution-ShareAlike 4.0 License. Selected authority identifiers and statements are retrieved from Wikidata under CC0; their references and qualifiers remain part of the verification path.