📚 rolling lexical diversity lab
Moving average type-token ratio calculator
Measure MATTR by sliding a fixed word window through your sample, averaging each window's type-token ratio, and showing how stable the vocabulary stays across the text.
| MATTR band | Fiction sample | Academic sample | How to read it |
|---|---|---|---|
| 0.40-0.49 | Repetitive dialogue or narrow scene | Formula-heavy or method-heavy text | Vocabulary recurs often inside each window. |
| 0.50-0.59 | Typical narrative draft | Typical essay section | Moderate diversity with stable topic focus. |
| 0.60-0.69 | Dense descriptive prose | Concept-rich literature review | Many new types appear as windows move. |
| 0.70+ | Compressed lyric or varied synopsis | High-density abstract or mixed corpus | Very high variety; check that token rules match the genre. |
| Sample length | Suggested window | Suggested step | Best use |
|---|---|---|---|
| 80-250 tokens | 50 tokens | 10 tokens | Flash fiction, abstracts, poems, short student answers. |
| 250-1,200 tokens | 100 tokens | 20 tokens | Essays, book reviews, scenes, article openings. |
| 1,200-5,000 tokens | 200 tokens | 50 tokens | Chapter sections, long reports, interview excerpts. |
| 5,000+ tokens | 500 tokens | 100 tokens | Corpus checks, full chapters, multi-document comparisons. |
| Token setting | Usually raises MATTR | Usually lowers MATTR | When to use |
|---|---|---|---|
| Case sensitivity | Keeping case-sensitive types | Lowercasing all tokens | Use lowercase for most manuscript comparisons. |
| Hyphenation | Counting compounds and parts | Joining compounds as one type | Use one rule across all drafts. |
| Numbers | Keeping every number as written | Replacing numbers with #NUM | Tag numbers for research prose with many figures. |
| Stopwords | Removing function words | Keeping all words | Keep stopwords for standard MATTR-style reporting. |
| Metric | Length sensitivity | Core calculation | Best comparison |
|---|---|---|---|
| Plain TTR | High | Total types divided by total tokens | Only texts with nearly identical length. |
| MATTR | Lower | Average TTR across moving windows | Drafts or excerpts with different lengths. |
| Root TTR | Medium | Types divided by square root of tokens | Quick normalization checks. |
| MTLD | Lower | Token runs until TTR falls below a threshold | Longer corpus-style lexical diversity analysis. |
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
To measure lexical diversity, use this type-token ratio moving average calculator to compare your excerpts. Type your text in, select your rolling window size, check how stable the MATTR is and then compare your drafts against each other.
Then there’s lexical diversity (which sounds like something only linguists would care about until you edit your manuscript and realize that you’ve got the same ten words on every page). That subtle re-use robs a story or an argument of its depth quicker then anything grammatical can. To get a more reliable indication of vocabulary variation (compared to the schoolbook version of type-token ratio), use moving average type-token ratio, or MATTR.
How to Use Moving Average Type-Token Ratio
Basically it takes a set window size and runs over your words, counting up how many unique word (types) is present within each chunk. Then it averages all those ratios. What you’ll see is whether your vocabulary are richly varied or constrained in consistent ways throughout. The method has one major advantage over simple word-counting: it solves the problem of type-token ratio’s decline as a passage grows in length. Because writers will tend to repeat themselves on longer works, raw TTR decreases even if they’re being creative. To get an accurate picture of how people use words, MATTR divides up the piece into overlapping windows, then it averages them together.
For example, a fresh-feeling short story might end up at 0.62; a tedious dialogue section may be closer to 0.48. But don’t take those numbers too literally; MATTR is useful only if you’re using the exact same rules for tokens and window sizes throughout your comparisons. If you change parameters, you aren’t comparing apples to apples anymore.
What about the windows? Think of a little window (say fifty tokens) that will react violently to any sudden burst of new words in the passage. Then imagine a large window (two hundred or more tokens), which smooths out the peaks and valleys and lets you see the wider flow of the text. And that’s important if you are comparing say a child’s story to a dense abstract from some academic article.
Once you select your step size and window size, the rest is handled by the calculator. But it all comes down to your decision. Select a single size setting throughout your project and the scores begins to speak rather than contradict each other.
Another subtlety are tokenization options. Do you want to preserve capitalization, or lowercase everything? Will you split up hyphenated compound words, or count them as one word? What about numbers. Do you want to leave them unchanged, or replace them with an identical placeholder for all? All these adjustments affect the number of unique types in the tokenized file; each adjustment bumps the overall average higher or lower.
Simple decisions, such as stripping out punctuation and leaving common function words unchanged, tend to help most manuscripts. Stopword removal makes every sentence look artificially diverse. It can increase (or decrease) your score by almost a tenth of a point, that’s enough to swing any given text into a different meaning group. Some writers think a high score means everything’s great. Not exactly. Throwing a lot of fresh images at readers every few words might result in a very high-scoring lyric poem. That doesn’t mean it’s immersive or controlled like a tightly focused scene that has a lower score.
The range between the highest and lowest scores is called the stability band. It shows you things you won’t see just by looking at the average. If the range is small (a narrow band), then there is consistent use of craft. If there are wild swings up and down, maybe the writer was leaning on dialogue tags or ran out of steam in certain places.
It’s one thing to see this kind of excerpt scored against something. But it’s another to see it compared with other things: a poem, an essay, a piece of fiction, some academic text. Each will have its own familiar distribution. And then you can compare: your poor-looking score for that literature review could turn out to be spot-on for a basic language practice exercise. You’re shown not just the raw number but where it falls within those genre norms. This helps you judge if the piece needs work or if it has been written in a way that fits the type of writing you are dealing with.
Nope. No sentences written by MATTR. But it will point out when your words is falling into familiar ruts and where they’re doing their job. See how the numbers line up? Change a setting here and there. Then the hidden structure of your writing begins to emerge. And that’s where the real craft develops: in this feedback loop between writer and measurement. It happens one rolling window at a time.

