📚 Content-word analyzer
Lexical density calculator
Measure content words divided by total counted words, tune the function-word list, and inspect POS-style buckets for drafts, excerpts, and study passages.
| Function group | Common examples | Why excluded from content count | Calculator handling |
|---|---|---|---|
| Determiners | the, a, this, those | They point to nouns rather than adding topic substance. | Always function words unless forced as content. |
| Pronouns | I, we, they, it, who | They replace nouns and often lower lexical density. | Included in every built-in base list. |
| Prepositions | of, in, with, between | They mark relationships between lexical items. | Expanded most in academic mode. |
| Auxiliaries | is, have, can, should | They carry grammar, tense, mood, or voice. | Can be forced as content for special analysis. |
| Conjunctions | and, but, because | They connect clauses or phrases. | Counted as function words by default. |
| Text type | Typical range | Reading feel | What to compare |
|---|---|---|---|
| Dialogue scene | 32-45% | Fast, personal, grammar-heavy | Compare speaker turns with narration separately. |
| Children's story | 38-48% | Concrete but repetitive | Check whether repeated function words dominate. |
| News report | 45-56% | Information-rich but direct | Compare lede and background paragraphs. |
| Academic abstract | 54-66% | Compressed, noun-heavy | Check if dense clusters need unpacking. |
| Technical manual | 50-64% | Procedural with terminology | Count numerals as content when data matters. |
| Category | Heuristic cues | Counted as content? | Useful caution |
|---|---|---|---|
| Nouns and names | Proper-case words, terms, unknown lexical candidates | Yes when noun toggle is on | Names can inflate density in bibliographic or legal text. |
| Lexical verbs | Action lists plus endings such as -ed and -ing | Yes when verb toggle is on | Auxiliary verbs stay function words unless forced. |
| Adjectives | Descriptive suffixes such as -al, -ive, -ous | Yes when adjective toggle is on | Some nouns share adjective-like endings. |
| Adverbs | -ly words and common manner markers | Yes when adverb toggle is on | Discourse adverbs may behave like connectors. |
| Numerals | Digits, dates, percentages, and measurements | Depends on numeral option | Use content mode for data-heavy passages. |
| Setting | Raises density when | Lowers density when | Best used for |
|---|---|---|---|
| Academic connectors expanded | Connectors are removed from content count | The denominator remains all tokens | Research abstracts and argument paragraphs. |
| Proper nouns as content | Names carry topic information | Reference lists have many names | Biography, history, news, literary analysis. |
| Numerals as content | Data points are meaningful terms | Dates are mostly formatting clutter | Technical prose, reports, and tables in prose. |
| Split contractions | Not and auxiliary parts are visible | Dialogue gains extra function tokens | Conversation, fiction, and readability checks. |
| Force content words | A word is lexical in context | A custom list is too broad | Domain words such as may, can, like, or shall. |
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
What percentage of a text is made up of content words relative to all words? This calculates lexical density in texts. It identify parts of speech, proper nouns, numbers, and function words. Function words is used to filter which words are included or excluded from the lexical density. If you write your own stuff, then editing with an eye towards lexical density turn into a useful tool.
How does it work? It is a measure of how many meaningful words there is compared to how many simply hold sentence together. Think of former as ‘content’ words (e.g., adverbs, adjectives, main verbs, nouns), and the latter as ‘function’ words (e.g., auxiliaries, prepositions, articles). The ratio of these two types of words determine how reading feels to the reader.
What Is Lexical Density?
The tool show you how close your own writing is to this spectrum when you paste in a sample paragraph. Thirty-five to forty-four percent tend to be light and conversational, good stuff for telling stories or having character talk. Go beyond sixty-five percent and the text begin to feel dense. It is great for an academic abstract or technical manual, but exhausting for typical general reader.
It doesn’t tell you whether this is right or wrong. It just shows you what the ratio is so you can see if it’s what you’re going for. And here’s why: It adjusts for genre. You want to know how much is in a childrens story? It is not very dense, since there are lots of easy connectors and repetition. What about a complex legal clause? Pretty high-density, as it loads up with qualifiers and specific nouns.
You pick the base list of functions word and everything turns on that. Academics will lengthen the list to include connectors that could otherwise creep into numerator. That’s what trips everyone up. Move from one setting to another and you’ll change density by eight or ten points, and not alter a word in actual piece.
There is another level of nuance involving proper nouns and numbers. On the one hand, a long list of proper nouns and dates are an inherently information-rich part of a news story. Sure, count those toward the content total. On the other hand, a novel may feature a similar series of names as stage-directions. If you separate out each name, they will drags down density.
Here’s where the tool comes into play. You can take same paragraph and run it through twice. Toggle on/off the proper-noun option. See which version pulls heavier. Those little choices are what explain why two drafts feel so different, even though they’re using about the same words.
Other controls include contractions and hyphenated compounds. Uncontract “can’t,” for instance, to “can not.” See where that hides the auxiliary? When you’re checking for readability, this might be helpful. Leave “well-known” intact as a single lexical unit. It’s more reflective of how people read it. None of these are a hidden switch. All reflect actual editorial decisions writers makes routinely. They often do this without even recognizing their impact on clarity and pace.
But the real juice happen in those side-by-side comparisons where you’re comparing several different draft of the same paper, all set to the exact same parameters. You can see how your student’s essay (which began at fifty-two percent) could of decreased to forty-one following their revision. What happened? They cut some filler phrases; they added some concrete verbs. The numbers aren’t as important than the change. When your ear is fatigued by reading the draft over and over again, they provides an objective anchor.
Cultural resonance is not captured by any calculator. Rhythm and tone are not. Awkward phrasing can’t be fixed by a beautifully balanced density score. But it does something nearly as valuable. The invisible skeleton of your sentences becomes visible. Instead of hoping that the prose will land where you want, you get to sculpt it deliberatley.
This means that lexical density is less about rules than about direction. A writer checks it, tacks accordingly, and stows it. She listen again to the passage, this time with different ears. The search isn’t for an ideal-percentage reading; it’s for a piece of writing that simply feels right for what it’s trying to do. Which may be dense, authoritative prose, or may be something lighter and more welcoming.
When you’re conscious of how it affects the reader experience, you begin to notice it all around, on the novels on your bookshelf, the email in your inbox, the directions on that coffee-packet, too. And that, in itself, sharpens everything you write.

