📊 Phrase repeat checker
Four-gram frequency counter
Paste a draft, excerpt, catalog note, or OCR page to find repeated four-word phrases, overlap density, and phrase echoes that can flatten prose.
| Repeat density band | Fiction or memoir | Academic prose | Blurb or query copy |
|---|---|---|---|
| Clean variation | 0% to 1.0% | 0% to 1.5% | 0% to 0.8% |
| Watch list | 1.1% to 3.0% | 1.6% to 4.0% | 0.9% to 2.5% |
| Patterned language | 3.1% to 6.0% | 4.1% to 7.0% | 2.6% to 5.0% |
| Heavy repetition | Above 6.0% | Above 7.0% | Above 5.0% |
| Text type | Useful control setup | What to inspect | Likely action |
|---|---|---|---|
| Novel chapter | Words only, overlap, lowercase | Unintentional narration echoes | Revise phrases repeated in nearby scenes |
| Dialogue scene | Keep contractions, sentence restart | Repeated speech habits | Keep voice markers, trim accidental loops |
| Back cover copy | Words only, strip marks, top 12 | Marketing phrase sameness | Replace weak repeated promise language |
| OCR cleanup | Words and digits, overlap, min 2 | Scanner artifacts and duplicated lines | Check repeated blocks against source pages |
| N-gram type | Sequence length | Best signal | Common blind spot |
|---|---|---|---|
| Unigram | 1 word | Keyword or filler frequency | Misses phrase repetition |
| Bigram | 2 words | Short collocations | Flags too many normal pairs |
| Trigram | 3 words | Phrase fragments | Can overreact to grammar frames |
| Four-gram | 4 words | Distinct repeated wording | Needs enough text for stable results |
| Normalization choice | Counts together | Keeps separate | Use when |
|---|---|---|---|
| Lowercase phrases | Phrase and phrase | Different punctuation only | You want editorial echo detection |
| Preserve exact case | Exact casing only | Phrase and phrase with capitals | You are checking typesetting or OCR |
| Boundary sentence marks | Words within sentences | Phrases crossing sentence ends | You want prose rhythm only |
| Content-word phrases | Core repeated terms | Function-word frames | You want topic motif checks |
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
There’s an odd blind spot that comes about if you spend long periods staring at a manuscript. The language fills in blanks. Your brain stops noticing repetition and what you read flows smoothly along. In fact, you may well be reading things many times, like passage “the morning light filtered through,” which occurs six times in forty pages.
It happens to everyone who works with words from novelists to technical writers. The weakness isn’t one of craft but of familiarity. If you’ve lived within a manuscript, you stop being able to see it from outside.
How to Find Repeated Words in Your Writing
Enter: mechanical checks. They is a mirror for those times when our own sight grows fuzzy. To do that, the four-gram frequency counter breaks up your sentence into chunks of four word and counts each chunk’s repetition rate.
Four words is a good number. It is long enough to be a clear phrase but short enough to catch accidental repetitions that goes beyond one paragraph or even a single sentence.
To use it, just paste in some text and hit Calculate. Here’s the calculator: Punctuation is removed from the text before processing, and all words are lowercased (i.e., “The morning light” is treated as the same thing as “the Morning Light”). That’s an important pre-processing step; without it, formatting differences like capitalization would hide the repetitions you’re looking for. You want to look at the pattern, not the casing.
The distinction between overlapping and non-overlapping scans matters more then most users realize. With an overlap scan, each and every potential four-word window gets checked, moving only one word at a time. That will pick up on exact echoes that may has been broken up by slight editing tweaks or offsets. When scanning for unintended narration loops, this is your move.
Stride mode does samples in non-overlapping blocks. While it can helps pick up on wider structural repetition, it won’t get hung up on single word changes. In short, it’s cleaner and less noisy. It helps tell if your writing has a rhythm that is almost hypnotic instead of varied.
The density score itself has to be interpreted within its context. Three-percent repeats may be great for a fiction draft where they serves some character voice or motif. They may be terrible for an academic abstract, which must be tightened because those repeats are redundant. The tool includes a reference table that spells this out clearly, so you know whether your text is in the heavy repetition zone or clean zone.
Zero percent repeat rate: don’t go chasing this dragon. You’d have prose that was so very different than it could come off as erratic or disjointed. What you want instead are spikes. Check if one particular four-gram is dominating your top list. Investigate. Is it a phrase you love, or something you’re relying on as a crutch when tired?
The other side of things involves OCR scans and digital archives. Repetition tends to be a mistake here, not something deliberately styled. Misalignment of sensors can causes lines or paragraphs to be duplicated in scanned pages. Four-grams are more likely to indicate a technical problem than an authorial flourish when we see them with some regularity in an OCR cleanup pass. Then it’s just a question of debugging. Marking up those chunks of text for human verification against the original document.
A lot of editing involves taking stuff out as much as putting things in. There’s repetition that you find that makes you cut away fat and sharpen the voice, because it makes you think: do I really need this sentence? Does it add anything other than echoing an idea already said? It’s not about making something perfect; it’s about becoming aware. When you’re aware of where the loops are, you can break them. And when you know that, you read differently. You begin to hear those echoes before they become typed words.

