📚 Lexical diversity lab
Corrected Type-Token Ratio Calculator
Measure CTTR as unique word types divided by the square root of two times the token count, with transparent token rules for fair text comparison.
Each preset uses original sample text and a realistic register so the CTTR result changes because of vocabulary variety, repetition, and preprocessing choices.
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
Corrected TTR is less sample-length sensitive than raw TTR, but the token rule still matters. Keep these settings identical when comparing drafts.
| Formula | Calculation | Length sensitivity | Best use |
|---|---|---|---|
| Raw TTR | Types / tokens | High; drops quickly as samples grow | Very similar text lengths only |
| Root TTR | Types / sqrt(tokens) | Moderate correction | Quick lexical variety check |
| Corrected TTR | Types / sqrt(2 x tokens) | Moderate correction with a two-token denominator | Book excerpts, abstracts, and comparable passages |
| MATTR | Mean TTR across moving windows | Depends on chosen window size | Long documents where local variety matters |
| Example at 500 tokens | Unique types | Corrected TTR | Reading signal |
|---|---|---|---|
| Controlled vocabulary page | 120 | 3.79 | High repetition or restricted word set |
| Balanced fiction excerpt | 180 | 5.69 | Moderate vocabulary spread |
| Varied nonfiction section | 240 | 7.59 | Broader terminology and phrasing |
| Dense academic abstract | 300 | 9.49 | Specialized terms and compact phrasing |
| Sample size | Reliability note | What to report | CTTR caution |
|---|---|---|---|
| Under 100 tokens | Very unstable | Use as a quick snapshot only | One unusual word can swing the score |
| 100-249 tokens | Rough comparison | Mention the short sample | Dialogue and lists can distort results |
| 250-999 tokens | Useful excerpt range | Report token rules and genre | Still compare like with like |
| 1,000+ tokens | Stronger passage profile | Consider section-level comparisons | Long documents may hide local shifts |
| Preprocessing choice | Usually raises CTTR | Usually lowers CTTR | When to use |
|---|---|---|---|
| Case-sensitive types | Names and sentence starts count separately | Not usually | Only when capitalization is meaningful |
| Light suffix grouping | Not usually | Running, runs, and run group closer | When word families matter more than forms |
| Stopword removal | Content words dominate the remaining set | Token count also falls | Topic-focused vocabulary comparisons |
| Hyphen kept as one token | Compound forms stay distinctive | Split parts are not counted separately | Technical, academic, and hyphen-heavy prose |
Want to check lexical diversity in your own writing? This is the corrected type-token ratio calculator. Enter a chunk of text, select your token rules, get back your type count, your token count, your CTTR, and the breakdown of formula behind it all. It’s an academic-sounding thing, lexical diversity, but it’s something that affects all writers and editors alike.
Here you are with two documents of roughly the same size. One is more interesting than the other. Why? Probably because there is more different words in it. That’s what corrected type-token ratio (CTTR) measures, and it gives you a score on one number. The key point is that it doesn’t have the primary weakness of simpler type-token ratio: longer things looks less varied by nature.
What Is Lexical Diversity?
The math is easy. Basically, we divide the number of unique words by the square root of two times the number of tokens (i.e., words). It means that you can compare a 900-word excerpt with a 300-word page without issues. Because the square root is a slow-growing function, you can rely on score regardless of sample size.
How accurate that number is comes down to how you set things up prior to calculating it. Do you count “ran,” “running,” and “run” as distinct words, or lump them into a category? Are words connected by tiny conjunctions like “and” counted separately, or should they be left joined as one word with hyphens intact? All of this determines what constitutes variety. Depending on these filters, the results for academic writing about special terminology could differ than those of a children’s writer repeating a word. A two-point difference in the final tally can mean all the difference.
People who is experienced with language will lock in their rules for word form and capitalization, to ensure they’re consistent when comparing samples. After locking in their rules, they can runs a draft through the calculator and see how it has changed over time. Often a draft feels baggy because the author is using all of those safe words. As the draft gets revised, new synonyms appear. Repeat words gets swapped out, which raises the score. Without the calculator, you’re left guessing whether or not your edits made the text better. With the calculator, you can see what has happened.
There’s no one number that paints the full picture. High CTTR might indicate precision or overuse of jargon. Low CTTR might signal deliberate simplification or a lack of skilful writing. What is right depends on context and genre. Spoken language (such as dialogue) typically has fewer unique words, which results in a lower score for these kinds of scene. Poetry appears simple, yet each word have significance. The tool cannot assess style, but knowing this helps avoid unfairly comparing two texts with vastly different styles.
The key is consistency and honesty. Note down what you choose, why, alongside the eventual number. Apply the same rules to each and it is then not only an oddity but a useful companion. A reminder that vocabulary can become stale and that a passage might be well stocked with words.
So what’s lexical diversity? It’s about noticing how words sound when they come out of your mouth (or onto a page). The corrected type-token ratio gives that texture a name, and allows you to track it in following drafts. As long as you have those numbers over time, you know whether or not you’re saying something new. You should of noticed the difference earlier.

