Unicode Normalization

Normalize pasted text to NFC, NFD, NFKC, or NFKD in the browser. Copy returns the normalized string only.

Last updated: October 6, 2026

Tool

This tool will load here.

What this tool does

This unicode normalization page takes pasted text and one form, then rewrites the text with the built-in String.prototype.normalize operation. The four forms are NFC, NFD, NFKC, and NFKD. NFC and NFD are the canonical forms. NFC prefers composed characters. NFD prefers decomposed characters. NFKC and NFKD are the compatibility forms. They can merge or split characters that are not canonically equivalent, so a ligature can become two letters and a compatibility letter can become the ordinary letter it decomposes to.

The result is that browser operation, not a table maintained on this site. The page shows the normalized string, whether it differs from the input, the code-point count before and after, and the UTF-8 byte count before and after. The byte counts use TextEncoder. Copy writes the normalized string only. It does not copy the counts.

The page does not render a per-character table. Listing each scalar is a different tool. Counting code points, UTF-16 units, and UTF-8 bytes without rewriting the string is also a different tool. Changing case is a different tool. This page answers a single question: what is this string in the form you chose?

NFKC and NFKD must not be read as a security filter. Compatibility decomposition is not a confusable check, and it is not a promise that two strings which look alike will become identical. The result note for those two forms says they are not a security canonicalization and not a confusable check. The page does not score spoofing, and it does not claim the output is safe to compare for security.

How to use

  1. Paste the text into the text area. Empty text is an error.
  2. Choose NFC, NFD, NFKC, or NFKD.
  3. Choose Normalize. Pressing Enter inside the text area does not run the conversion. Enter in the form list does. Copy writes the normalized string. Clear empties the text and the result and leaves the selected form as it is.

If any lone surrogate is present, Normalize reports an error and clears the output. The value is not passed to normalize. If the text is longer than 200,000 UTF-16 code units, the output is cleared and no partial string is shown. A well-formed surrogate pair, such as an emoji that uses two UTF-16 units, is one code point and is allowed.

Example

The composed letter é, U+00E9, is already NFC. Choosing NFC reports that the string did not change, with one code point before and one code point after. Choosing NFD produces e followed by the combining acute accent U+0301. The changed line says yes. The code-point count goes from 1 to 2. The UTF-8 byte count changes with the new sequence. The same split is what String.prototype.normalize returns for NFD in this browser.

The ligature fi, U+FB01, stays a ligature under NFC because the decomposition is a compatibility decomposition. Choosing NFKC produces the two letters fi. The code-point count goes from 1 to 2, and the result includes the compatibility warning. The angstrom sign U+212B is likewise a compatibility equivalent of Å, U+00C5. NFKC rewrites it to U+00C5. NFC follows the canonical form of that character and does not apply the compatibility mapping.

A grinning face emoji that is one code point stays one code point under NFC, and its UTF-8 length is 4 bytes. A string that contains a script tag is ordinary text. Normalization does not parse it as markup, and the output shows the same characters. A lone high surrogate or a lone low surrogate is rejected. Two hundred thousand a characters are accepted. One more unit is rejected and the previous result is cleared.

Limits

  • One form per run. Cap of 200,000 UTF-16 code units. No partial result.
  • A textarea turns CR and CRLF into LF before Normalize runs, so a bare CR will not survive this field.
  • No character-name database, no grapheme segmentation, no spoofing verdict, and no network request.

The text stays in the page. It is not stored. The counts describe the string before and after the selected form. They are not a second inspector and they are not a claim about how a font will draw the glyphs.

On this page

Related Articles

Related guides

Related solutions

Related Tools

Need another tool ?

Open the free tools — no signup.

FAQs

newsletter signup

Lorem ipsum dolor sit amet, consectetur adipiscing elit.
Innovative Solutions For Modern Needs
Copyright © 2026 Yallasolve. all rights reserved.