Online Tool Store Online Tool Store
📚 Education

· 4 min read

How to Build a Vocabulary List by Frequency

Heshan Fernando

Co-founder & COO

Heshan Fernando is the Co-founder and Chief Operating Officer of Ceyentra Technologies, where he leads project management, engineering, and research and development strategy. With over nine years of industry experience, he is passionate about transforming complex customer challenges into practical, high-impact solutions. His customer-centric leadership has enabled multidisciplinary teams to consistently deliver secure, scalable, and industry-grade digital products that create lasting business value. View on LinkedIn

Share

How to Build a Vocabulary List by Frequency

Learning vocabulary alphabetically is obviously inefficient and people do it anyway, because a dictionary is sorted that way. Learning by theme is better and still misses the point.

Learning by frequency is dramatically more efficient, and the reason is a property of language itself.

Why frequency ordering works so well

Word frequencies in any natural text follow a steep distribution. A small number of words appear constantly; most words appear once or twice.

The pattern is close to Zipf’s law: the nth most common word appears roughly in proportion to 1/n. The most common word is about twice as frequent as the second, three times the third, and so on.

The practical consequence is coverage. In English, the most common 100 words typically account for around half of all running text. The first 1,000 cover a large majority. Beyond a few thousand, each additional word adds progressively less.

That’s why frequency-ordered learning delivers comprehension so much faster than any other order. The first few hundred words are doing enormous work, and you meet them constantly, which means they’re also the easiest to retain.

It’s also why the curve flattens. Going from 3,000 to 4,000 words adds far less coverage than going from 500 to 1,500, and at that point reading widely beats drilling lists.

Words knownApproximate coverageWhat it enables
100~50% of running textRecognising structure
1,000Large majorityGist of simple text
3,000Higher stillReading with a dictionary
BeyondDiminishing per wordRead widely instead

Stopwords: remove or keep

Remove them for vocabulary learning. The most frequent words in any English text are “the”, “of”, “and”, “to” — function words a learner acquires immediately and doesn’t need on a list. Stripping them surfaces the content words that actually carry meaning.

Keep them for stylistic analysis. Function word frequencies are one of the more reliable signals in authorship attribution, precisely because writers use them unconsciously and consistently. For that purpose they’re the data, not the noise.

Choosing a source text

The list describes the text you fed it. That’s obvious and easy to forget.

A single novel gives you that author’s vocabulary. A collection of news articles gives you journalistic register. Technical documentation gives you terminology.

For general language learning, a broad mixed corpus is better than any single source. For learning to read a specific thing — a subject’s literature, an author, a genre — build the list from exactly that, because the specialist vocabulary that matters won’t appear in a general list.

Size matters too. A single article is too small for frequencies to be meaningful; a few hundred thousand words starts being representative.

Grouping inflections

Group them for vocabulary work — “run”, “runs”, “running” and “ran” are one item to learn, and counting them separately makes a common verb look like four uncommon ones.

Keep them separate for stylistic analysis, where the choice between forms is part of what’s being measured.

Common mistakes to avoid

  • Building a list from one short text and treating it as representative of the language.
  • Keeping stopwords in a learning list, where the top twenty entries are words you already know.
  • Counting inflections separately, which fragments common verbs.
  • Learning past the point where coverage flattens, instead of reading.
  • Ignoring the words appearing exactly once — in a large corpus they’re the long tail, and there are always more of them than you’d expect.

How to do it with Frequency Word List Builder

The Frequency Word List Builder reports coverage alongside counts.

  1. Paste a text large enough to be representative of what you want to read.
  2. Remove stopwords for vocabulary work; keep them for stylistic analysis.
  3. Group inflections if you’re building a learning list.
  4. Read the coverage figure — it tells you when more words stop paying off.

Other language tools are in the tools directory.

Frequently asked questions

Why learn vocabulary by frequency?

Because coverage rises steeply at first. The most common few hundred words appear constantly, so learning them delivers comprehension faster than alphabetical or thematic order.

Should inflections be grouped?

For vocabulary learning, yes — “run”, “runs” and “running” are one item. For stylistic analysis, keep them separate, since the choice between forms is part of what you’re measuring.

How much text do I need?

Enough to be representative of what you want to read. A single novel gives you that author’s vocabulary; a mixed corpus gives you something closer to the language.

Final thought

Read the coverage number, not the word count. It tells you the one thing a frequency list is for: when to stop learning lists and start reading.

Try the free Frequency Word List Builder

#frequency-word-list#vocabulary-learning#corpus-analysis#word-coverage#online-tools#free-tools