📚 Lexical diversity lab
MTLD calculator
Measure of Textual Lexical Diversity estimates how long a text maintains a high type-token ratio before vocabulary repetition pulls it below a threshold.
| MTLD band | Approximate score | Typical sample | Reading signal |
|---|---|---|---|
| Low diversity | 20 to 40 | Highly repeated classroom, dialogue, or controlled-language text | Many words recur before the TTR threshold is reached. |
| Moderate diversity | 40 to 70 | General prose, summaries, shorter reviews, or accessible nonfiction | Vocabulary varies but common wording remains prominent. |
| High diversity | 70 to 100 | Essays, analytical reviews, feature writing, and mature fiction passages | Longer lexical runs appear before repetition lowers TTR. |
| Very high diversity | 100+ | Academic abstracts, literary prose, dense exposition, or specialized analysis | Many unique types are introduced across extended token spans. |
| Text type | Useful sample size | Expected MTLD tendency | Comparison caution |
|---|---|---|---|
| Children's narrative | 80 to 250 tokens | Often lower because repeated character and action words are intentional | Compare by age level and passage length. |
| Dialogue or transcript | 150 to 500 tokens | Moderate to low because pronouns and discourse markers repeat | Speaker turns can reduce lexical density. |
| Book review | 150 to 700 tokens | Moderate to high depending on plot, craft, and evaluation vocabulary | Repeated title or author names may depress MTLD. |
| Academic abstract | 120 to 300 tokens | Often high because discipline terms and nominal phrases are dense | Short abstracts can inflate uniqueness. |
| Technical documentation | 120 to 600 tokens | Mixed because repeated nouns improve clarity but lower diversity | Do not treat repetition as poor writing automatically. |
| Literary prose | 200 to 1000 tokens | Often high when imagery, setting, and action words vary | Scene focus can raise or lower scores. |
| Metric | What it counts | Length sensitivity | Best use |
|---|---|---|---|
| MTLD | Average token run before TTR drops below a threshold | Lower than simple TTR for longer texts | Comparing lexical diversity across prose samples. |
| Forward MTLD | Factors from first token to last token | Can reflect opening repetition or novelty | Checking how a passage begins lexically. |
| Reverse MTLD | Factors from last token to first token | Can offset end-heavy vocabulary shifts | Balancing order effects in short excerpts. |
| Type-token ratio | Unique word types divided by total word tokens | High sensitivity to sample length | Quick vocabulary variety context only. |
| Lexical density | Content words compared with all tokens | Depends on tagging and function-word rules | Studying informational compression, not diversity alone. |
| Control | Default | Effect on MTLD | When to change it |
|---|---|---|---|
| Threshold | 0.72 | Higher thresholds usually create shorter factors and lower scores | Use non-default thresholds only for planned comparisons. |
| Lowercasing | On | Merges Book and book into one type | Turn off for proper-name or capitalization studies. |
| Numbers | Remove | Prevents dates and figures from inflating type counts | Keep numbers for technical or quantitative prose analysis. |
| Hyphenated words | Keep as one | Treats well-known and long-term as single lexical types | Split for studies that need separate root words. |
| Stopwords | Keep all | Preserves the standard MTLD behavior over running tokens | Filter only for exploratory content-word comparisons. |
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
MTLD: Measure your lexical diversity by averaging forward and reverse factor scoring. This MTLD calculator include all calculations, such as tokens and customizable token rules. It also allows for comparison across text registers and more.
What does “lexical diversity” mean? It’s an academic word for something that becomes surprisingly practical once you start editing your own writing… Be it a manuscript or an essay, or review someone else’s. This matters until you edit your own writing, be it a manuscript or an essay, or review someone else’s. Is the writing fresh and new, or is it circling a couple of word so many times that the reader wants to tear her hair out?
How to Use MTLD for Better Writing
That question has a name: the MTLD (Measure of Textual Lexical Diversity), which measure the diversity of your vocabulary. Unlike some other ways of measuring diversity, the MTLD holds up against actual variations in text lengths in the real world.
MTLD: To understand MTLD better, consider how it works: instead of just counting unique words against total words, it tracks how far a text can travel before its running type-token ratio drop below a chosen threshold. Instead of counting the number of unique words compared to total words, it looks at distance: How far does it go before it dips under some preset threshold? In other words, think of it as the average distance a writer gets in a piece before they feel pressured into repeating themselves. So a higher MTLD is more stretched out, not using same word so much.
That one idea shifts how you approach reading what you’ve written. People tend to think if vocabulary is richer, the MTLD will also be higher; and yet, that’s not true at all. By design, kids’ books repeat important words over and over so young readers can follow along. An academic abstract may load up on specific nouns used within a discipline. Though harder to digest, it’ll end up scoring more higher. You pick the register for comparison, and it helps make sense of where the band lies for the genre. A spoken conversation might have a low score, but that is common in writing.
As for setting the right starting point (the threshold), beginners underestimate its importance. The classic setting sits at 0.72. Increasing it makes everything end sooner and lowers total score. Dropping it means that factors will stay open a bit longer allowing text to meander further before closing down. For most uses of writing, keeping close to the default is best. Once you get very far away from the default, you’re no longer comparing apples to apples.
And there’s another twist: running it in reverse tells us when the end starts introducing new words versus recycling what came before. It also tells us when new words start appearing on the opening end (running forward). Taking an average of both sides evens out order-related oddities found with brief passages. All three are run by the tool, so you can find out if your piece accumulates vocabulary gradually, or saves some for the surprise reveal.
What does it mean when I say that small decisions matter? The decision on token rules matters. Do you want to treat “well-known” as one lexical item or two? What about numbers? Should they stay or go? Function words are common and thus can inflate your token count but don’t bring much new variation to the table. Does their presence matter? You can set up the analysis so that these decisions reflect what you’re really trying to capture and thus aren’t cosmetic at all. If your goal was to compare texts with exact definitions, such as in a technical manual where repetition is used for clarity, it shouldn’t count against you. This is different than a marketing brochure that keeps repeating itself.
MTLD is not a quality score. Repetition doesn’t always signal sloppiness; sometimes it’s used for emphasis or rhythm. Moderate diversity isn’t uncommon in clear, effective prose. The metric flags an unintentional slip into monotony. If you see the score sink low relative to the usual band for the register you’ve set yourself, go hunting for spots where a new synonym could improve things.
Length matters, but much less than raw type-token ratio. Under fifty tokens, the results is noisy. A hundred or more gives the factors enough room to form properly and produce a steadier number. This explains why comparisons typically use similar-sized excerpts from two or more docs instead of, say, comparing a whole short story to one paragraph.
So what does all this do for you? It gives you a disciplined process to point out what your own ear already hears when you read something. When language becomes stale, your ear picks up on it. And now, you have a measurement to quantify it and take action based off an informed response rather than guessing. After running through a couple of samples, you begin to see patterns emerge within your writing. Some things pull your score upward; other thing pull it downward. It’s a quiet yet helpful feedback loop.
If that’s the case the next time your draft starts to feel flat and you don’t know how to fix it, plug it into the calculator. It won’t rewrite the sentence for you, but it’ll give you an answer (the number) that points to where the vocabulary is beginning to take hold. And from there, you can make clear revision decisions. Every good analytical tool should of be able to do this: take a vague impression and shape it into something you can work with.

