There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
Having said that, the categories that shrank all did so by a big enough percentage to also shrink in absolute number of words, so at least that isn't a problem.
It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
I used the examples of Latin to Spanish and English, or Old English to new English.
I say that to say, languages changing over time is to some people not actually something they believe happens. The facts are right there in front of you, but some people cannot have their minds changed no matter how much sense you make.
I'm not going to finger-scroll or down-arrow the whole thing.
Don't break scrolling. Please.
This is something I've been thinking a lot about. We have trended from subjective language to objective language. Why?
Computing. Software is written with objective language. Everything is clearly unambiguously defined. Blue is no longer a category, it's #0000FF. Logic must always reduce to a binary truth value. Most of what we have to talk about is somehow relative to software. Software even structures most of what we write! We don't just talk to each other, we tweet, email, message, post, search, etc. These structures each imply a specific set of phrase structures that can make sense.
Lately, it's hard to go even a day without reading some complaint that such and such was written by "AI" (an LLM). Why is this so obvious? Well, the core advantage that LLMs provide is that they don't compute. Inside an LLM, there is no arithmetic, no logical branches, no truth values. Phrases aren't generated to define or to resolve. They are generated to continue. Sure, we can direct the story to follow the steps of logical deduction, but that isn't anything like calculation. An LLM simply isn't invested in logic, precision, correctness, etc. the way we expect modern writers to be. It's not the em-dashes or the word choice that illustrates this, it's the fundamental perspective of the system.
We are sorely missing subjectivity. Natural language never was, and never will be, computable. You can't reduce a natural story to binary truth values without choosing an arbitrary perspective that resolves its ambiguity. The more precisely abstract our language gets, the more detached from reality our stories become. The more objective our assertions about reality are, the less relevant they can be.
My answer to this is to make the arbitrary choice of perspective a first-class feature. If we can explicitly decide what meaning is relevant, we should be able to weakly solve natural language processing. It seems like a pretty simple and obvious idea, but so far is easier said than done.
I was pretty confident in my English but I couldn't understand half of the sentences since I didn't know any of the vocabulary related to cooking. From tools (spatula) to ingredients (parsley) or even actions (frying). My delusions of proficiency shattered as I found out my everyday vocabulary is smaller than an elementary schooler.
What I see is shifting from a focus on individuals and close personal relationships to a focus on the functioning of society and, well, identity.
It seems like a swing of the pendulum to me. Encouraging people to create their own identities, creating vocabulary to describe one's identity, and advocating for a restructuring around society that is relatively collectivized economically, but culturally celebrates individual diversity. I'm not a professional historian, but I think that's somewhat of a new phenomenon historically.
And it's clearly a reaction to the perception that the ideals of equality and being uniformly kind to each other we're not enough to bring about a just and equitable society, and if anything were abused in order to suppress diversity. We still have to be kind to each other at a basic level, but it's not enough anymore.
So maybe now there's too much focus on that stuff and not enough focus on the part where you should actually just be nice to your neighbors and be humble and all that. Or, maybe people are disproportionately writing about it, because it's a new thing, and a lot of people know a lot to say about it, and a lot of people want to talk about it.
I don't think it makes a shred of sense to conclude that somehow the world is now more equal than it was in the 50s. It's disappointingly far from equal, though, and I think the change in lexicon is a change in people's focus specifically because the old system clearly wasn't working, and people are eager to try something different.
Linguists have resorted to spying on people in parks and shops and restaurants, to get "natural" language unperturbed by observer effects, audience tailoring, the filter of writing, etc. (In most cases that would not be considered ethical today!)
The gold standard is something like: get acculturated to the language community you wish to study, make some acquaintances and friends, somehow finagle an invite to their homes or workplaces or social venues, and then with their permission, record their conversations as unobtrusively as you can.
There is so little data on how language is really used it is quite bedevilling. As soon as people know you are listening they change how they speak.
For reasons unknown to me, instead of using the word with the meaning they probably are trying to convey, they rather use vague analogies or metaphors.
It always made me question if Americans in contrast to the British are a bit afraid to be precise.
1. TV vocabulary is 2,000 words
2. High school vocabulary is 10,000 words
3. College vocabulary is 30,000 words
4. English language has 1,000,000 words
This is good news - you can be proficient in a foreign language by learning only 2,000 words!
I live and work in a foreign country, and have a modest but functional grasp of the language here. Which is to say, I know how to say 'sustainability', 'union-negotiated collective agreement', and 'offensive conduct in the workplace'.... but if you asked me the words for 'cedar', 'robin', 'pond', or 'linen' I would be struck dumb. And yet presumably every ten year old I walk past in the street would know those words as comfortably as I knew them in English as a ten year old in England.
Why? Having spent a good amount of time living in Shanghai, I found it important to be able to understand menus. But there's no pressure to know the words for supermarket items; you can just go to the supermarket and look for the item.
Otherwise your point is correct; all semantic words are equally difficult and which ones you know depends on the things you like to talk about. Grammatical words are more difficult, and more important, but this is so widely understood that language-learning material already treats them as an entirely separate class of things to learn.
> Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed.
(1) You seem to want COCA, which includes a bunch of transcribed telephone calls.
(2) Word frequencies are still the wrong concept. If you want to understand a particular document, you need to understand almost all of the words that appear in that document. (You'll be able to learn some of them from their use in the document.) If you decide to learn a list of "frequent" words, you're unlikely to be able to understand more than a couple of isolated sentences in any given document.
Language has to do with our Monkeysphere, that humans over long periods of time in the past were limited to a very small subset of people they interacted with socially and closely. A few institutions likely had a large effect on your language like the church depending on where you were. After that it was the people you interacted with to stay alive. Because almost everything was in person or person to person transfer of information a lot of socially encoded clues were involved which lessened the need for well defined words.
Books were the first stage of homogenizing language as they could be shared over long distances and to many people, but more sequentially than latter forms of communication. After that radio and TV had a huge effect, for example the 'General American' used in broadcasts that was based heavily on a midwestern accent.
As we encroached on the 70s and 80s the previous technological advances and things like high speed interstates and trucking shrank America to something you could drive across in less than a week, and you could reach anywhere by voice nearly instantly. Suddenly people in California, Texas and New York all could be in the same meetings and local colloquialisms would need explained, so people would trend to a shared vocabulary.
It's also odd to me to say an LLM isn't subjective. Each LLM has it's own behavior, it's that there are like 20 or 30 big LLMs in all, and people are using them millions to billions of time so we're getting that one LLMs language everywhere. And that's why I disagree and will say natural language is computable, but it's also lossy and probabilistic. And for the most part it's single prompt and being ran by the user for the cheapest price possible.
In 1953 people were not exposed as much to different groups of people far away.
> (Formerly in Israel, when a man went to inquire of God, he said, “Come, let us go to the seer”; for he who is now called a prophet was formerly called a seer.)
I love that one because it translates so well into English because we also have an older and a newer word for close enough to the same thing. Though when saying older and newer I should note that the words “prophet” and “seer” have been used consistently since the very first English translation, Wycliffe’s in ~1382. Maybe it’s more the feel of the words. (In Bible Gateway’s catalogue of English translations, only the CEV doesn’t use both of those words: <https://www.biblegateway.com/verse/en/1sam9.9>; I disqualify OJB as not English.)
That’s the only one I can think of off the top of my head that’s specifically about language, but there are plenty of others about changed customs over time, in Old and New Testaments.
Most of these kind of website have the animation on the centre of the screen. This one animate at the bottom
https://news.ycombinator.com/pool
But not in this case. It's a slow Sunday afternoon and getting a few upvotes quickly is enough
Must just be the combination of that increase as well as other words decreasing in usage I suppose. E.g. perhaps we're a bit less keen, but also much less passionate, so keen ends up making the cut.
Of course, American corporate-speak is unique from normal conversations, and people intentionally use vagueness to make sure there's general agreement before saying something concretely that might be contentious or cause problems.
Just earlier I was talking to them, trying to say "when life gives you lemons" and being unable to find an equivalent phrase. I hope they understood my lemons reference regardless.
I didn't learn English by reading vocabulary lists or dictionaries. Pretty much all was learned by context.
There are many reasons to suspect that information theory applies to human language. And information theory tells us that the rarest symbols are the least predictable and therefore carry the most information.
That seems not to be available at either the Internet Archive nor Archive Today:
<https://web.archive.org/web/20200000000000*/https://ello.co/...>
<https://archive.is/https%3A%2F%2Fello.co%2Fdredmorbius%2Fpos...>
Sic transit gloria data.
Copying lyrics from my favorite songs of the 80s was helpful. And then later on TV series and movies in english.
I've watched Friends, HIMYM and Californication* more than once (it's probably the reason I swear so much) and those really helped.
Now I raised my kid in english: from the beginning she was only ever allowed to watch TV shows in english. And she's always been to international schools, in english too.
The other day she taught me a simply word which, surprisingly enough, I didn't know: a "coaster" (a drink coaster).
My absolute favorite is when she uses english grammar and applies it to french: the rules about where adjectives do go vary between french and english and she regularly messes it up in french (supposedly her "native" language but at this point I wonder whether english ain't a first language, and french second).
And one thing she has that I'll never have is a beautiful, proper, english accent. Mine shall forever be a strong french one.
Saying "I'm keen to do x" or "he's keen to do y" is a fairly common normal usage but having said that, wouldn't normally be my first choice of language as a millennial, maybe does sound a bit old fashioned
Oughtn't really affect emotive language like 'keen', though.
Look through the words scattered around this page. Every one of them is among the most commonly used words in the English language.
They come from a 2023 list of about 2,800 words, shown to cover over 90% of general English use, intended for people learning the language.
This 2023 list is an update of an earlier one, made in 1953, which identified about 2,300 words as the essential vocabulary to everyday life at the time.
Between the two lists, 70 years apart, about 600 words were dropped, and over 1,100 were added. The rest remained as is.
Some of the changes make immediate sense: Telegraph dropped out; computer was added, along with website and blog. Tobacco was replaced by cigarette. Motherhood became mom, and dad was added too, though fatherhood was never on the list to begin with. The world changed and vocabulary surely followed.
But also: apple didn’t make the new list. Neither did fork, soap, umbrella or leaf, for example. It’s not that these things vanished from everyday life, but many hands-on words became less central to the core vocabulary. Dog stayed; goat and donkey didn’t. Bread stayed; breadmaking ingredients—flour and wheat—dropped. Cook is on the new list. Boil, bake, and fry are not.
And many of the words that were added, such as mortgage, corporation, appropriate, analysis, fairly, and despite, don’t look anything like the ones that were discarded. In fact, they are mostly abstract concepts that don’t look like anything at all.
A
Added to 2023 list
A
Removed from 1953 list
A
In both lists
How the Words We Teach English Language Learners Changed, and What That Says About Us
These “essential vocabulary” lists are called the General Service List (1953) and the New General Service List (2013, revised in 2023).
They were designed as teaching tools for people learning English as a second language, built from real-world usage data and extensively tested. The aim was a vocabulary list with as few words and as much coverage of everyday English usage. That coverage is high, over 90% for the 2023 list1 and about 84% for the 1953 list.2 To account for that much of the language, they had to track a significant portion of whatever people were actually reading and saying. A word earned its place by appearing often enough, across enough contexts, to be hard to avoid for the average person in an English-speaking society. Open the word lists panel on the right to browse through all the entries in both lists.
While these were practical tools, built to capture which words people need most, the answer also doubles as a snapshot of ordinary life, 70 years apart: what people were expected to engage with, and had to deal with, in their daily lives.
Treating the lists as an indirect record of day-to-day life, I went through the differences between them from a few angles: what the words were about, how tangible were they, and to which parts of speech they belonged.
I started by using a common linguistics research tool3 that sorted all of the words by meaning. It assigned each word to one of 21 subject categories, based on typical usage: Food and Farming, the Body and the Self, Government and Public, Language and Communication, and so on. Together, they suggested what kind of world each list was built for.
Data source: all words run through the UCREL Semantic Analysis System (USAS).
The categories that shrank are mostly those that have to do with the immediate, physical world, while the gains are those furthest from it.
To better see this, the categories can be oriented spatially. Some categories describe you: your body and emotions. Some describe what’s immediately around you: food, objects, and the natural world. Others name systems you participate in: government, institutions, society, and culture. And some have no location at all: abstract concepts, reasoning, and processes. When we look at the data grouped by physical scope, the pattern of change seems to point in a specific direction.
Data source: all words run through the UCREL Semantic Analysis System (USAS), then grouped into five umbrella-term domains. Expand to view the full category breakdown.
In hindsight, the shift makes sense. By 1957, four years after the original list was published, white-collar workers outnumbered blue-collar for the first time in US history. And by 2000, fewer than one in four workers did manual labor. The stuff people encountered in their daily lives, what they needed to talk about, and the systems they had to navigate all changed.
The vocabulary lists, built from the language of their respective eras, tracked those changes. The shifts reflect a life that moved further from its own making: less tied to tools, animals, food, and the body; more tied to national or global institutions, categories, systems, and ideas.
The new vocabulary—mortgage, legislation, perspective, involvement, improvement, assumption, evaluation—are words you can’t weigh, point to, or hold in your hand. These words shape your life, but they do it without occupying physical space.
To better understand whether there was a shift in vocabulary describing the physical world, I compared each word to a database that rates its tangibility on a scale of 1 to 5 (called a “concreteness rating4”). A rating of 5 means you can experience the word directly with your senses, while a rating of 1 means you can’t.
The highly concrete end dropped the most: from 21% of the total in 1953 down to 14% in 2023.
Kernel density estimation, bandwidth 0.08 · Data source: Brysbaert et al. (2014)
This shift matters because abstract and concrete words are processed by our brains in different ways. When you read axe, your brain doesn’t just decode letters, it reaches for something: an image, a weight, or a gesture. The word activates both a verbal label and a sensory trace. Psychologists call this dual coding:5 Concrete words travel through two channels, verbal and sensory; abstract words travel through one. Two channels mean two retrieval pathways, which is why concrete words are “stickier,” easier to hold in a line of thought and faster to recall. Abstract words, on the other hand, are purely verbal, and have to be understood through language alone.
To put it another way: Concrete words are easier for us to process because they are bundled with a web of associations, tactile experiences, and memories that anchor their meaning. Here’s a more detailed view of the shift away from concrete language:
Data source: Brysbaert Concreteness ratings for 40 thousand generally known English word lemmas. The dataset provided 99.8% coverage of the words in both lists. 6 words not in the dataset are not included in this chart: as, dialog, english, gaiety, madden, old-fashioned.
While concrete words have sensory grounding to carry their meaning, abstract words rely on other parts of speech to specify, soften, or sharpen what they mean. That help tends to come from one particular corner of the language: adverbs.
The 2023 list contains more words overall (2,809 vs. 2,284). All changes mentioned in the text reflect each category's share of its list, not raw counts. Data source: NLTK (Natural Language Toolkit), with manual correction of mislabeled words.
Axe doesn’t need an adverb to modify it; you know what it is. But acceptable, relevant, and adequate come with conditions, qualifications, and degrees that need to be spelled out. Adverbs do precisely that: language to calibrate language.
Look through all the adverbs that were added:
absolutelyactuallyalongsideanymoreapparentlyapproximatelyautomaticallybadlybarelybasicallybrieflycarefullycertainlyclearlycloselycompletelyconsequentlyconstantlycurrentlydeeplydefinitelydifferentlydirectlydramaticallyeasilyeffectivelyentirelyequallyeventuallyexactlyextremelyfairlyfaithfullyfinallyfirmlyfirstlyforeverfrequentlyfullyfurthermoregenerallygentlygraduallygreatlyheavilyhencehighlyhopefullyimmediatelyincreasinglyinitiallylargelyliterallymainlymerelymonthlymostlynaturallynearbynearlynecessarilyneverthelessnewlynormallyobviouslyoccasionallyoriginallyoverseasparticularlypartlyperfectlypersonallypossiblypotentiallypreciselypresumablypreviouslyprimarilyprobablyproperlyquicklyquietlyrapidlyrarelyreallyreasonablyrecentlyregardlessregularlyrelativelyrespectivelyroughlysecondlyseriouslysignificantlysimilarlysimplyslightlyslowlysomewhatspecificallystronglysubsequentlysuccessfullysuddenlysurelysurprisinglytotallytrulytwicetypicallyultimatelyunfortunatelyusuallyvirtuallywidely
Most of the adverbs specify degree, frequency, certainty, and extent. Some hedge (somewhat, partly, relatively, possibly, approximately). Others assert (absolutely, definitely, entirely, exactly, precisely). They're all doing the same kind of work: Take a statement and tell you how much of it is true, how often, and how certain. It’s as if the world now requires you to be more precise about everything.
Bread survived both lists. Flour, wheat, harvest and bake didn’t. The word for what sustains us remained essential, while the words for how we’d make it weren’t. That might be the most honest summary of what happened.
The world that made the 2023 list is more regulated, more connected, and in many ways more capable than the one behind the 1953 list. It’s a world further than our kitchen or home, reaching across economies, institutions, and democracies.
Today's vocabulary reflects a life that isn’t self-contained, but rather more systemic. It’s not so much about what’s within arm’s reach, but more about the larger world we navigate through. That sort of long-distance connection requires a particular kind of language: expansive, abstract, and precise. And language, it turns out, can’t help itself. It keeps track.
I compared two prominent vocabulary lists for English learners: the General Service List (GSL, 1953; 2,284 words) and the New General Service List (NGSL 1.2, 2023; 2,809 words). I labeled words appearing on both lists as “remained” (1,656), words only on the 1953 list as “removed” (628), and words only on the 2023 list as “added” (1,153).
The GSL words came from the Simple English Wiktionary GSL. The NGSL words came from the official NGSL 1.2 file (“alphabetized and lemmatized for research”). The NGSL uses lemmas (one entry per word family); the GSL sometimes lists inflected forms as separate headwords. These lists track word forms deemed worth teaching based on frequency and usefulness, not abstract concepts. For example, the word being is on the GSL as its own headword and was not included in the NGSL, But this doesn’t mean the concept of existence left the language. In the NGSL it falls under be, which is on both lists.
Why treat the lists as a portrait of everyday English?
Both lists were built for teaching, but external research suggests each covers a large share of everyday language use. About 84% of general English for the GSL and about 90% for the NGSL, depending on the text and how words are counted. I did not re-run those corpus analyses myself. I take the published materials as given and rely on coverage figures from the list authors and from follow-up studies (including an independent check on American English by Stoeckel, 2019 – see footnotes). That’s why the differences between the lists felt worth examining as more than a curriculum update, with the caveat that neither list is a neutral census of culture.
Where can I find the data?
All tagged words: remained, removed, and added, with semantic tags, concreteness ratings, and part-of-speech labels are in this public spreadsheet. You can also browse the word lists in the word-list panel on the right side of this page.
How did I sort words by meaning?
Each word was tagged with the UCREL Semantic Analysis System (USAS), using the 21 top-level categories. USAS also assigns much finer sub-categories (there are hundreds), but I stayed at the top level so the charts could show broad shifts without splitting the words into overly granular bins.
I chose not to correct mislabels. For example, hammer, nail, and wax are all tagged “General and Abstract Terms”, but in USAS’s finer tags, they read as actions (“to hammer,” “to nail,” “to wax”), not objects. Out of context, many of these words go multiple ways, and it didn’t feel right to override that case by case; USAS is an established linguistic framework, and swapping in my own judgment would mix two different standards.
For the second chart, I grouped USAS’s 21 categories into five “scope” domains (self, local, institutional, social, abstract). That grouping is my editorial choice, not part of USAS. It came from noticing a spatial quality to the trends seen across the 21 categories.
How did I measure concreteness?
Concreteness ratings come from Brysbaert et al. (2014). I used these as-is. Six words weren’t in the database and were left out of the concreteness charts: as, dialog, english, gaiety, madden, and old-fashioned.
How did I tag parts of speech?
Parts of speech were tagged with NLTK, simplified to five categories. Here I did intervene, but only when a word was clearly mislabeled (132 words, 3.8% of the list; mostly adjectives mislabeled as nouns). When a word can act as more than one part of speech depending on context, I left the tag as-is and deferred to NLTK as the established framework, using the primary, most common label. Here I also used an LLM strictly to help flag potential errors in NLTK’s output.
What are the limitations?
Both lists rank words by frequency, then apply learner-focused curation. Michael West’s 1953 list especially reflects period pedagogy. He favored general-purpose vocabulary over emotional or highly specific words, not just whatever appeared most often (Therova, 2020, summarizing West, 1953, pp. ix–x). Some of what looks like “1950s life” may also be how mid-century ESL teaching filtered the language. I still treat the lists as a portrait of everyday English because both cover a large share of running text and speech that goes well beyond the classroom.
Footnotes