Input

Canonical composition — the usual choice for text on the web and in databases.

Normalized output

Start typing above — the results update as you write.

Convert text between NFC, NFD, NFKC and NFKD, strip zero-width and directional characters, and see exactly which code points changed and how many invisible marks the text was carrying.

What is Unicode Text Normalizer?

Two strings can look identical on screen and still be different at the character level, which is why a search misses a name, a duplicate check fails, or a form rejects a caption that appears to be well inside its limit. The Unicode Text Normalizer converts text between the four standard normalization forms, optionally removes invisible characters, and reports what changed — in characters, code points and per-character categories.

Unicode normalization is the process of replacing a string with one of its standard equivalents. Canonical equivalence covers sequences that mean the same character: e sharp can be stored as a single code point or as an e followed by a combining acute accent, and both render the same way. Compatibility equivalence is broader and folds characters that are merely similar in use, such as a ligature, a full-width Latin letter or a superscript digit into their plain forms.

The four forms follow from that. NFC composes to the shortest canonical form and is the usual choice for web content, databases and identifiers. NFD decomposes, so accents become separate combining marks — the shape older macOS file names used, and still common in paths exported from that platform. NFKC composes and also applies compatibility folding, which is useful when normalizing user input for search. NFKD decomposes and folds, which is the most aggressive reduction and should be reserved for matching rather than storage, because it discards distinctions some documents care about.

The stability row answers a question worth asking before changing anything: which forms is your text already stable in? If NFC is listed, running the NFC conversion will leave the text byte-for-byte identical, and you can see that instead of assuming it.

Invisible characters are the second half of the job. Zero-width spaces, zero-width joiners, soft hyphens, a byte order mark at the start of a file, Arabic letter marks and bidi controls are all characters that occupy no visible space but change length, matching and layout. The tool lists every one it finds with its offset, its hex value and a plain-language name, and the optional removal step reports how many were deleted. This is the fastest way to explain a caption that a platform rejected at a count you could not see, or a search term that will not match the text on the page.

Everything runs locally. Normalization is a call to the platform's own string function, the invisible-character scan is a character-by-character classification, and the download button gives you the converted text as a plain file. Two side notes the tool does not hide: canonical normalization never changes how text looks, while compatibility folding (NFKC, NFKD) can, so test with your own content before applying it in bulk.

Log files, exports and pasted quotations are where the tool earns its keep. Running a sample through it to remove invisible characters takes one click, and the reported count tells you whether a file is clean or merely looks clean. Convert a whole set with the same form and the text stops shifting between applications — and whenever a normalizer run changes something you care about, the before-and-after counts and the stability row show it before you commit to the result.

How to use Unicode Text Normalizer

  1. Paste the text you want to inspect or convert.
  2. Read the stability row to see which forms the input already satisfies.
  3. Choose a target form: NFC for web and database storage, NFKC or NFKD when you need aggressive matching.
  4. Tick the invisible-character box if the output is going into a search index, a duplicate check or a length-limited field.
  5. Check the before-and-after counts, then copy or download the normalized text.

When to use Unicode Text Normalizer vs related tools

Use it whenever text crosses a boundary: form input into a database, exported file names, copy pasted from a word processor, translated content, or anything that has to match exactly in a search or a duplicate check. The Character Count Tool shows the length consequences in more detail, including UTF-8 bytes and user-perceived characters.

When the text behaves strangely in layout rather than in matching, the Bidirectional Text Tester identifies directional controls and shows what each one does, which is a narrower and more informative view than a full strip.

Privacy & Security

This tool runs entirely in your browser — no data ever leaves your device. There is no server round-trip, no upload, no logging, and no account required. Your input is processed locally using client-side JavaScript and is never stored, transmitted, or accessible to anyone else. When you close the tab, everything disappears.

Frequently asked questions about Unicode Text Normalizer