🔤 Unicode text tool
Diacritic counter
Count accent marks, combining characters, affected letters, and Unicode normalization changes in pasted text.
| Rank | Base | Diacritic | Count | Share |
|---|---|---|---|---|
| No diacritics counted yet. | ||||
| Density band | Marked letters | Typical source | Review signal |
|---|---|---|---|
| Light | 1-3% | Names and titles | Spot check |
| Moderate | 4-10% | Mixed-language notes | Check consistency |
| High | 11-25% | Accent-rich prose | Normalize carefully |
| Dense | 26%+ | Vietnamese or scholarly text | Preserve marks |
| Mark class | Examples | Common use | Filter name |
|---|---|---|---|
| Tone marks | ́ ̀ ̉ ̃ ̣ | Stress or tone | Tone marks |
| Diaeresis and ring | ̈ ̊ | Vowel quality | Diaeresis |
| Cedilla and ogonek | ̧ ̨ | Consonant or nasal marks | Cedilla |
| Length marks | ̄ ̆ ̂ | Length or shape | Macron |
| Unicode form | What it does | Diacritic effect | Best use |
|---|---|---|---|
| NFC | Composes text | May hide marks | Storage |
| NFD | Decomposes text | Exposes marks | Counting |
| NFKD | Compatibility split | Expands more forms | Cleanup |
| Stripped | Removes marks | Plain base letters | Search keys |
| Sample type | Expected marks | Risk | Check |
|---|---|---|---|
| Bibliography | Names | Lost spelling | Compare NFC |
| OCR scan | Mixed | Random marks | Sort by order |
| Transliteration | Macrons | Flat vowels | Filter length |
| Tone language | Stacked marks | Meaning shift | Count clusters |
DISCLOSURE: This post may contain affiliate links, meaning when you click the links and make a purchase, I receive a commission. As an Amazon Associate I earn from qualifying purchases.
Then one day they is gone. And then suddenly you see them: those little markings, above or below letters. Removing every accent from a name on a résumé makes it harder to read, making it look bland and anonymous instead of being a unique part of someones identity. Or maybe you work at an archive digitizing an old collection of manuscript, and all the cedillas have been replaced by some arbitrary question mark? The Optical Character Recognition Software got confused.
Small symbols can make a big difference. They establishes meaning in tonal languages. They protect how words sound in European written languages. They ensure accuracy of scholarly references. To lose them isnt just to misformat text, it’s to lose information.
Why Small Symbols Matter
That’s where the tool comes in (above). It breaks down unicode strings into their individual letters and marks so that you gets a sense of just how complex your text might be, or even if you’ve lost some. When you see something like é, most folks view that as one character. But on a computer, it could be two: two code points that happen to be glued together. For searching and cleaning up data, that difference are important. Normalize to NFD, and suddenly, the mark detaches from the letter. Once you can count them, the counter will suggests doing so in that format for your analysis.
The structure is there, exposed; standard viewers might overlook it. How heavy is your text marked? This falls on a spectrum, and our reference table on the page breaks it out into bands (light, medium, dense). Light tends to mean proper nouns or titles with accents sprinkled throughout. High density suggest a text that relies on diacritics for grammatical function, such as Vietnamese or specialized linguistic transcripts. You should of know where your text lies so you can determine how aggressively to process it. A document with a lot of information may not handle heavy text processing without losing important meaning. On the other hand, you may be more inclined to reduce for maximum compatibility if the doc is lighter. The calculator allow you to visualize this tradeoff before committing to any particular normalization path.
Sometimes people think all accents are the same, but they’re not. Accents is used for different things. A diaeresis mark means that you’re saying each of the two neighboring vowels as a separate unit; a tone mark completely changes the value of the syllable it applies to. A length mark tells how long a vowel lasts. Using the filters to select just one or the other lets you zero in on a class of errors. For example, if you’re dealing with some Sanskrit text being transliterated into English, then you probably do want to be looking at macrons. Maybe you’re looking at French poetry, acute and grave accents may be more important to look at there. By filtering for class, you can audit for individual kinds of errors. It’s going from this sort of wide overview, down to an individual investigation.
But first, let’s look at the comparison grid which will give you an instant sanity check on whether or not your data is intact. You’ll see number of code points as well as the number of normalized code points. A big difference between those two numbers mean you’ve got a mixture of precomposed characters that were split into their constituent parts. That’s not necessarily bad, just different; it changes how the text sorts and searches.
The stripped out length represents a kind of baseline. How does that compare against the marked one? The difference between them tell you how much the diacritics actualy represent. Sometimes: nothing. Other times: quite a lot. But that is all true. Diacritics arent just decorative; they’re a necessary part of the language infrastructure.
If you’re fixing up a spreadsheet gone bad, if you’re sifting through user generated content, or if you’re archiving old documents, knowing what characters make up your text is the start of ensuring it stays accurate. The tools give you a view of the invisible. And then once you can actualy count those things, you can determine whether you want to preserve them, remove them or work around them. That’s where the value lies, that ability to make something visible when before it was a kind of hazy notion of “messy text.

