📊 Honoré’s Statistic Calculator
Measure vocabulary richness using Honoré’s formula — enter token, type, and hapax counts or paste your text
| Text Type | Typical R Score | V1/V Ratio | Notes |
|---|---|---|---|
| Everyday Conversation | 50 – 200 | 0.60 – 0.80 | High repetition of common words |
| Children’s Literature | 100 – 300 | 0.50 – 0.70 | Simple, repeated vocabulary |
| News / Journalism | 300 – 600 | 0.40 – 0.60 | Moderate lexical diversity |
| Academic Prose | 500 – 900 | 0.30 – 0.50 | Technical terms boost uniqueness |
| Literary Fiction | 700 – 1200 | 0.20 – 0.40 | Rich descriptive language |
| Poetry | 800 – 1500 | 0.10 – 0.30 | Dense, unique vocabulary |
| Legal Documents | 200 – 500 | 0.35 – 0.55 | Formulaic but technical |
| Metric | Formula | Scale Dependent? | Best For |
|---|---|---|---|
| Type-Token Ratio (TTR) | V / N | Yes (decreases with N) | Short, equal-length texts |
| Honoré’s R | 100 × log(N) / (1 − V1/V) | Less so (log corrects scale) | Varying text lengths |
| Guiraud’s Index | V / √N | Partially | Medium-length texts |
| Herdan’s C | log(V) / log(N) | Low | Cross-language comparison |
| Yule’s K | 10⁴ × (ΣVr×r² − N) / N² | Low | Large corpora |
| Token Count (N) | Min Recommended V | Expected V1 Range | R Stability |
|---|---|---|---|
| 50 – 100 | 25+ | 15 – 60 | Low — high variance |
| 100 – 300 | 50+ | 40 – 180 | Moderate stability |
| 300 – 1000 | 100+ | 80 – 600 | Good stability |
| 1000 – 5000 | 200+ | 150 – 2500 | High stability |
| 5000+ | 500+ | 400+ | Very reliable |
| Register | V1/V Ratio | V1/N Ratio | Interpretation |
|---|---|---|---|
| Spoken Informal | 0.65 – 0.85 | 0.25 – 0.45 | High repetition of filler words |
| Written Informal | 0.55 – 0.75 | 0.20 – 0.40 | Moderate lexical variety |
| Written Formal | 0.40 – 0.60 | 0.15 – 0.30 | Balanced lexical use |
| Scientific | 0.30 – 0.55 | 0.12 – 0.25 | Precise repeated terminology |
| Literary | 0.20 – 0.45 | 0.10 – 0.20 | Rich, non-repetitive vocabulary |
Honoré’s free Statistic calculator can be used to figure out the richness of a given vocabulary within seconds. Simply paste a block of text into it for automatic calculation, or enter number of tokens, types, and words that appear only once you’ve counted yourself.
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
If you’ve ever read something that seems to have surprising and varied vocabulary, then congratulations: it was probably pretty lexically rich. In other words, it had high Honoré’s Statistic. That’s the name of one number that measures lexical richness. The measure, devised by French linguist Pierre Honoré in the late 1970s, go beyond basic word-counting methods and zeroes in on unique words that occur just once in a given piece of text. Known as words that appear only once, these uncommon words signal if an author is working within their comfort zone or trying out new linguistic territory.
What is Honoré’s Statistic?
Because longer texts tend to have more word repetitions, they are “corrected” for document length. That’s why raw word counts here might be misleading. To get a fair score, Honoré uses a formula that divides the natural log of total tokens by (1 minus percent of Hapax words). This lets you reasonably compare two 2,000-word short story or a 200-word blog post. It’s still used by scholars to analyze student essays and presidential speeches.
In practical terms, the input represents the following: Total tokens = the total number of words you tallied (including repetitions) Types = the distinct word forms after ignoring capitalization/punctuation Hapax legomena = words that the author used a single time A large percentage of one-time-only words boost the overall score, indicating few repetitions among total words. This is our definition of a rich vocabulary.
To compare, the calculator also provides related metrics such as Guiraud’s index and the type-token ratio. As texts becomes longer, the simple type-token ratio decreases, which is a problem for researchers who work with data of different lengths. Honoré’s formula prevents this slippage. It would of been helpful, then, if you’re using samples of different lengths.
But a metric is never the complete story. A document full of technical language may have high Honoré’s Statistic score, yet feel stiff to read. A text that has a low one may still be clear for its target audience. Variety isn’t necessarily quality; the metric only indicates variety. The number of tokens also plays an important role. You need to have more than a hundred tokens before things settle down; one random word can make your hapax ratio shoot up and down like crazy. So use at least a hundred tokens, which is what the tool advises. The logarithmic correction doesn’t make comparing text on the same word count unnecessary either.
Many writers who try running their draft through the calculator are surprised. They realize a scene that seemed vivid was supported by overuse of their favorite word. They also realiszed that an academic tone scored better than expected, since it uses jargon that has its own form. These realizations don’t always occur naturaly. You have to run the numbers and force yourself to notice what’s hidden in plain sight.
For reference, you can compare your scores against other registers. Conversation generally scores poorly because we use a limited selection of words efficiently. Literary fiction rates up because writers mix things up in order to hold the reader’s attention. Poetry frequently soars right to the top by jamming particularity into relatively few words. This is all tendency different than prescription. It will give you an idea of how your own work lines up on that scale.
So in the end, I think of Honoré’s Statistic less as a judge and more based off a mirror. Does this sound like something people will read? Is it convincing? That’s not what Honoré’s Statistic can say. But it can show you which language tools you’re most comfortabley with using. When you get a clear picture, you decide whether or not you want to change.

