Lexical Diversity Calculator

Your text

Diversity

Start typing above — the results update as you write.

The Lexical Diversity Calculator reports several length-corrected diversity measures at once — plain type-token ratio, the Guiraud and Herdan indices, and moving-average ratio at four window widths — so no single bias decides the answer.

What is Lexical Diversity Calculator?

Every lexical diversity measure corrects for something, and each correction introduces its own distortion. The Lexical Diversity Calculator shows several at once and a table of the moving-average ratio across four window widths, on the principle that a number you can compare with its alternatives is more useful than one presented alone.

The vocabulary is worth fixing first. A token is a running word as it appears: the word the twice counts as two tokens. A type is a distinct word: the same two occurrences count as one type. The plain type-token ratio is types divided by tokens, and its flaw is well known. It falls as a text grows, because the same common words keep recurring, so it cannot compare a short sample with a long one.

The Guiraud index divides types by the square root of the token count, and the Herdan index divides the logarithm of types by the logarithm of tokens. Both damp the length effect in different ways and both are reported here with the raw ratio, so you can see how much the correction moves the figure.

The moving-average type-token ratio is the length-stable measure favoured in modern work. It slides a fixed-size window across the text, computes types over tokens inside each window, and averages the result. Because every window is the same size, the figure does not drift with document length. The table shows the value at windows of 50, 100, 200 and 500 tokens, which turns the choice of window from a hidden assumption into something you can read off. A window wider than the text collapses to the plain ratio, and the page notes that rather than letting the collapse look like a real measurement.

No single figure is labelled correct. A high value means the text keeps introducing new words; a low one means it recycles a small set. Neither is good or bad without a genre. Poetry and advertising copy often score high; legal and scientific writing often score low because a controlled vocabulary requires exact repetition of defined terms. The tool reports measurements and the length caveat, and leaves the judgement where it belongs.

Two practical notes. Diversity measures need a few hundred tokens before they settle, so a short sample is flagged rather than trusted. And the tool counts surface forms, not lemmas, so run and running are two types; if you need lemma-based diversity, normalise the text first. For most editing purposes the surface measure is the honest one, because a reader sees the forms, not the dictionary entries.

The letters MATTR stand for moving-average type-token ratio, and the measure exists because vocabulary diversity has to be compared fairly across texts of different lengths. A window of fifty is sensitive to local repetition, while a window of five hundred approaches the whole-text figure. Reading the table from fifty to five hundred shows how much the window choice matters for your particular text, which is information a single number cannot give. If the value falls steeply as the window widens, the vocabulary is spread thinly; if it holds steady, new words are arriving at a consistent rate.

How to use Lexical Diversity Calculator

  1. Paste a draft of at least a few hundred tokens; the measures need length to stabilise.
  2. Read the moving-average ratio at the reference window, then the plain type-token ratio and the two indices.
  3. Check the tokens and types rows, which are the counts every ratio is built from.
  4. Compare the four window widths in the table to see how sensitive the measure is to window size.
  5. Read the warning if the sample is short, and treat the figures as provisional until it is longer.

When to use Lexical Diversity Calculator vs related tools

Use the lexical diversity calculator when you want several length-corrected measures and a window-sensitivity check in one place. When the question is richer vocabulary overall and you want the moving-average ratio as the headline with the classic ratios beneath it, the Vocabulary Richness Scorer presents it that way.

If the low diversity is caused by a handful of terms dominating the text, the Keyword Density Analyzer shows which ones, and the Overused Word Finder lists the words that recur most, including those returning within a short distance.

Privacy & Security

This tool runs entirely in your browser — no data ever leaves your device. There is no server round-trip, no upload, no logging, and no account required. Your input is processed locally using client-side JavaScript and is never stored, transmitted, or accessible to anyone else. When you close the tab, everything disappears.

Frequently asked questions about Lexical Diversity Calculator