Vocabulary Richness Scorer

Your text

Richness

Start typing above — the results update as you write.

The Vocabulary Richness Scorer measures how varied your word choice is, using a length-stable moving-average type-token ratio as the headline figure and the classic whole-text ratios beside it for comparison.

What is Vocabulary Richness Scorer?

Vocabulary richness is the simplest idea in text analysis and the one most often measured badly. Count the distinct words, divide by the total, and you have a type-token ratio — a number that falls mechanically as a text grows, because common words repeat. The Vocabulary Richness Scorer puts the length-stable measure first and the misleading one second, with the reason stated.

The headline figure is the moving-average type-token ratio, or MATTR. It slides a window of a fixed size across the text, computes the ratio of distinct to total words inside each window, and averages the results. Because every window is the same length, the figure does not drift downward simply because the document got longer, which makes it the only one of these measures that can compare a paragraph with an essay.

Below it sit the classic ratios. The whole-text type-token ratio is reported because it is what most people mean by richness, along with the Guiraud index, which divides distinct words by the square root of the total, and the share of words used exactly once. That last figure, the hapax share, is a useful second signal: a draft with a high proportion of words appearing a single time is reaching for vocabulary rather than recycling it.

The window size is adjustable, and the default of one hundred words is a reasonable compromise. A small window measures local variation and can be noisy on short texts; a large window approaches the whole-text ratio it is meant to correct. If your text is shorter than the window, the measure collapses to the plain ratio, which the page notes rather than disguising.

No band or pass-fail is attached, and that restraint is the point. Richness is genre-dependent. Legal writing repeats defined terms deliberately, scientific writing uses a controlled vocabulary on purpose, and literature repeats words for rhythm. A low ratio can mean thin vocabulary or it can mean correct terminology. The tool reports the measurement and the length caveat; the writer judges.

The practical reading is comparative. Measure a draft, revise a paragraph that felt flat, and measure again. Because the moving-average ratio is stable across lengths and the underlying counts are visible — words, distinct words, sentence count and syllables per word — the change you see is a real change in word choice rather than an artefact of how much you pasted.

Lexical richness is measured in more than one way for a reason, and the measures disagree on purpose. A short poem can score a high ratio simply because it is short, while a long novel with an enormous working vocabulary scores lower, which is why the moving-average figure is the one to compare across lengths. The hapax share and the Guiraud index are reported as second opinions rather than as alternatives to choose between: when all three agree, the reading is solid, and when they diverge the divergence is itself the finding, usually because text length is doing more of the work than word choice. Treat the panel as a set of corroborating measurements rather than as a leaderboard.

How to use Vocabulary Richness Scorer

  1. Paste a draft of a few hundred words; richness measures need length before they settle.
  2. Read the moving-average type-token ratio first — it is the length-stable headline figure.
  3. Compare it with the whole-text ratio and the Guiraud index in the table below.
  4. Adjust the window size if you want a local rather than a document-wide reading.
  5. Revise a flat paragraph, paste it back, and compare the two readings like for like.

When to use Vocabulary Richness Scorer vs related tools

Use the richness scorer when you want to know how varied the word choice is overall. When the question is specifically how the ratios behave across different window sizes, the Lexical Diversity Calculator tabulates MATTR at four widths alongside the Herdan index.

If the draft's problem is not thin vocabulary but too much concentration on a few terms, the Keyword Density Analyzer shows which words are taking up the most room, and the Overused Word Finder lists the words the text leans on hardest.

Privacy & Security

This tool runs entirely in your browser — no data ever leaves your device. There is no server round-trip, no upload, no logging, and no account required. Your input is processed locally using client-side JavaScript and is never stored, transmitted, or accessible to anyone else. When you close the tab, everything disappears.

Frequently asked questions about Vocabulary Richness Scorer