https://www.english-corpora.org/glowbe/ The corpus of Global Web-based English (GloWbE; pronounced "globe") is unique in the way that it allows you to carry out comparisons between different varieties of English. GloWbE is related to many other corpora of English that we have created, which offer unparalleled insight into variation in English. GloWbE contains about 1.9 billion words of text from twenty different countries. This makes it about 100 times as large as other corpora like the International Corpus of English, and it allows for many types of searches that would not be possible otherwise. In addition to this online interface, you can also download full-text data from the corpus. Click on any of the links in the search form to the left for context-sensitive help. You might pay special attention to the comparisons between countries and virtual corpora, which allow you to create personalized collections of texts related to a particular area of interest. == https://en.wikipedia.org/wiki/Corpus_of_Contemporary_American_English https://www.english-corpora.org/coca/ The Corpus of Contemporary American English (COCA) is the only large, genre-balanced corpus of American English. COCA is probably the most widely-used corpus of English, and it is related to many other corpora of English that we have created, which offer unparalleled insight into variation in English. The corpus contains more than one billion words of text (20 million words each year 1990-2019) from eight genres: spoken, fiction, popular magazines, newspapers, academic texts, and (with the update in March 2020): TV and Movies subtitles, blogs, and other web pages. Click on any of the links in the search form to the left for context-sensitive help, and to see the range of queries that the corpus offers. There are three main ways to search the corpus: First, you can browse a frequency list of the top 60,000 words in the corpus, including searches by word form, part of speech, ranges in the 60,000 word list, and even by pronunciation. This should be particularly useful for language learners and teachers. Second, you can search by individual word, and see collocates, topics, clusters, websites, concordance lines, and related words for each of these words. Note that some of these searches are unique to COCA and iWeb. Third, you can search for phrases and strings. And because the corpus is optimized for speed, searches for substrings (*ism, un*able) and phrases are very fast, e.g.: got VERB-ed, BUY * ADJ NOUN, "gorgeous" NOUN -- and even high frequency phrases like: from ADJ to ADJ, phrasal verbs, or NOUN NOUN. You might pay special attention to the comparisons between genres and years and virtual corpora, which allow you to create personalized collections of texts related to a particular area of interest. ==== English-Corpora.org Description (corpora arranged by size) iWeb 14 billion words from the Web NOW 9.73 billion, Web news, 2010-last month GloWbE 1.9 billion, Web, 20 countries Wikipedia 1.9 billion, Wikipedia Hansard 1.6 billion, British Parliament COCA 1.0 billion, US, 1990-2015 EEBO 755 million, 1470s-1690s COHA 400 million, US, 1810s-2000s TV 325 million words in 75,000 very informal TV shows Movies 200 million words in 25,000 very informal movies Supreme Court 130 million words, 1790s-present TIME 100 million, US, 1923-2006 SOAP 100 million, US, 1990s-2000s BNC 100 million, British, 1980s-1993 CORE 50 million words, Web genres Strathy 50 million, Canada