📖 Vocabulary load lab
Unfamiliar word density calculator
Paste a passage, choose a known-word baseline, add your own course vocabulary, and measure how much of the text may feel unfamiliar to a target reader.
The calculator marks a word as unfamiliar when it is outside the selected known-word list, outside your custom known words, and not softened by the proper-name or short-word rules.
| Density band | Unfamiliar rate | Typical reader effect | Revision response |
|---|---|---|---|
| Light | 0% to 4% | Smooth reading with few stops | Usually no glossary needed |
| Steady | 5% to 8% | Noticeable new vocabulary | Add context around key terms |
| Heavy | 9% to 14% | Frequent pause points | Preview terms before reading |
| Dense | 15% or more | Specialist or close-study load | Glossary and scaffolding help |
| Known-word baseline | Best fit | Words included | Use case |
|---|---|---|---|
| Starter reader | Early readers | High-frequency basic words | Decodable and simple passages |
| Elementary core | Upper elementary | Daily school and story words | Classroom reading checks |
| Grade 8 general | General readers | Broad common prose words | Books, essays, and articles |
| College common | Adult readers | More abstract everyday terms | Nonfiction and criticism |
| Filter choice | What it changes | When to use it | Watch for |
|---|---|---|---|
| Proper-name soften | Reduces name penalty | Fiction and history passages | Capitalized sentence starters |
| Content denominator | Ignores common function words | Vocabulary-focused review | Higher density percentages |
| Unique repeat mode | Counts each hard term once | Glossary planning | Misses repeated friction |
| Lite suffix cleanup | Groups related word forms | Fast draft comparison | Rough word-family estimates |
| Passage type | Expected pattern | Useful target | What to inspect |
|---|---|---|---|
| Picture book | Few unfamiliar words | 0% to 4% | Rare words and names |
| Middle grade chapter | Some challenge words | 4% to 8% | Repeated story terms |
| Study guide | Unit vocabulary clusters | 6% to 12% | Terms students know already |
| Academic abstract | Dense specialist phrasing | 12% or more | Jargon concentration |
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
What do educators call the friction in a text? They call it unfamiliar word density: how many new words a passage throws at a reader in one sitting. To put it another way, it’s an exact measurement of unfamiliarity in form of a percentage.
So what does that mean? Because higher the number, the less likely a student is to breeze right through, struggle but cope, or simply give up. Why? The theory behind it is simple: if the vocab barrier become too great, reading comprehension fails. And research shows that beyond two to five percent unfamiliar words in a given text, decoding slows down and understanding begins to slip away.
How to Measure Hard Words in Text
With the tool, you can figure out that breaking point beforehand. You paste in your passage, select your baseline (which words does the reader already know?), and system flags all the rest.
It’s not as simple as saying ‘the more hard words, the worse.’ It’s also about knowing which words is hard for which people. Mitochondria might be second nature to a biology major but completely alien to a third grader. This distinction is handled by the baseline selection, which allows you to specify what your audience already knows.
A lot of writers assume that they’re familiar with the same things as their readers. A word might appear frequently in daily news, so you think it must be common, but to a young reader, or an English language learner, it may be entirely new. That’s where custom glossary terms comes into play. Words that you’ve previously taught, such as metaphor or foreshadowing, shouldn’t of be contributing to your density score. Those aren’t the unknowns anymore. Adding them as known gives you a clearer picture of what real problem areas might be in this specific text. Because the calculator accounts for this, it outputs its results based off only on the new terms, making that number a lot more actionable when it comes to planning your lesson.
A second nuance is how you treat names and proper nouns. Every character name is going to appear as an unfamiliar word at first glance in piece of fiction, but it will repeat after a while. Now if you counted each word (Hamlet, Elizabeth) as a hard term, your density score would shoot through the roof, falsely. Instead most people want to ignore/soften probable names to get at words in the text that is the content vocabulary not just labels. That leaves the signal clean.
On the page there’s a reference table that classifies bands from light to dense. If your load is light then the text is probably accessible for independent reading; if your load is heavy, then you’ll have to provide support material or pre-teach terms before having students read it independently. Knowing what band your text is on helps you know how to assign it (as homework?) and/or read it aloud with guidance.
So it’s not necessarily about making the number smaller. Sometimes you’re pushing advanced readers with high-density text, and sometimes you have some jargon that you want to sneak into the text in a planned way. The point is intentionality. You should understand what the numbers mean. Fifteen percent unfamiliar words in your academic abstract may well be fine if it’s a graduate seminar, but a disaster if it’s an eighth-grade science class. The right answer depends on context.
And by tweaking the settings around suffix variants, or around plural forms, you can get a sense of how much it moves as you include those sorts of things. It lets you test the text against different reader profiles. This means vocabulary load measures empathy. To measure it means stepping outside your fluency in order to look at the text from the perspective of a person who is building a vocabularly. You notice where they’re getting stuck, then choose between simplifying the writing for them or supporting the learning process for them. And either way, you’re removing the blind spots that turn reading into something more like work than discovery.

