Unicode Normalization: NFC, NFD, NFKC and NFKD Explained

Unicode normalization forms compared: what NFC, NFD, NFKC and NFKD each do, which one to use for storage and search, and why NFKC destroys styled text.

Concepts8 min read
The letter e with an acute accent shown as one code point and as two

Unicode lets you write the same character more than one way. รฉ can be a single code point, U+00E9, or two โ€” e at U+0065 followed by COMBINING ACUTE ACCENT at U+0301. Both are correct, both look identical, and neither is equal to the other in a byte comparison.

Normalization is the process of converting text to one canonical form so that comparisons work. There are four forms, and choosing the wrong one silently destroys data.

Why multiple representations exist

Historical compatibility, mostly. When Unicode absorbed the existing national character sets of the 1980s and 1990s, it took their precomposed characters with them so round-tripping would be lossless. Latin-1 had a single character for รฉ, so Unicode has one. But Unicode also has the general mechanism of base plus combining mark, which can produce the same result.

So there are now two encodings of รฉ, and there will be forever โ€” Unicodeโ€™s stability policy forbids removing either.

The general case is worse than two. Vietnamese แบฟ can be written four ways. Korean Hangul syllables can be one precomposed character or a sequence of conjoining jamo. Any text arriving from an unknown source may be in any of them.

The two axes

The four forms are two decisions crossed with each other.

Composition. Do you want the precomposed single character (composed, C) or the base-plus-marks sequence (decomposed, D)?

Equivalence. Do you apply only canonical equivalence โ€” sequences that genuinely are the same character โ€” or also compatibility equivalence (K), which folds characters that mean the same thing but look different?

CanonicalCompatibility
ComposedNFCNFKC
DecomposedNFDNFKD

The rules are defined in Unicode Standard Annex #15, which is the authoritative document.

Canonical equivalence: NFC and NFD

Canonical equivalence means the sequences are the same character, differently written. Converting between them is lossless in both directions.

InputNFCNFD
รฉ (U+00E9)U+00E9U+0065 U+0301
e + U+0301U+00E9U+0065 U+0301
ร… (U+00C5)U+00C5U+0041 U+030A
ร… (U+212B, angstrom sign)U+00C5U+0041 U+030A

That last row is worth a look. U+212B ANGSTROM SIGN is canonically equivalent to ร… โ€” the standard says they are the same character โ€” so normalising in either form collapses it. This is the correct behaviour, and it surprises people.

NFC is the default choice. It is what the W3C recommends for the web, what most text arrives in already, and what you want in a database. NFD is mainly useful when you need to strip diacritics: decompose, then drop everything in the combining marks range, and cafรฉ becomes cafe.

macOS is the notable oddity โ€” its file system stores names in a variant of NFD, which is why a filename copied from a Mac can compare unequal to the same filename typed on Linux.

Compatibility equivalence: NFKC and NFKD

Compatibility equivalence is much more aggressive. It folds characters that are semantically the same but visually distinct, and it is lossy. You cannot undo it.

InputAfter NFKCWhat was lost
๏ฌ (U+FB01 ligature)fithe ligature
ยฒ (superscript two)2the superscript
โ…ฃ (Roman numeral four)IVthe single-character numeral
ใŽ (square kg)kgthe compact form
๏ผจ๏ฝ…๏ฝŒ๏ฝŒ๏ฝ (fullwidth)Hellothe fullwidth forms
๐—›๐—ฒ๐—น๐—น๐—ผ (bold sans)Helloall the styling
แดดแต‰หกหกแต’ (superscript)Hellothe superscript forms
โ’ฝโ“”โ“›โ“›โ“ž (circled)(H)(e)(l)(l)(o)the circles

That table is the single most important thing on this page for anyone using styled text.

Why NFKC destroys styled text

Every style on this site is a mapping into a Unicode block whose characters carry a compatibility decomposition back to plain ASCII. That is not a flaw in the styles; it is what the standard says those characters are. U+1D5DB MATHEMATICAL SANS-SERIF BOLD CAPITAL H is formally defined as a compatibility variant of H, because in a formula it is an H, printed differently.

So any system that applies NFKC to incoming text will turn ๐—›๐—ฒ๐—น๐—น๐—ผ into Hello and ๐’ฎ๐’ถ๐“‡๐’ถ๐’ฝ into Sarah. Not corrupt it โ€” convert it, correctly and deliberately, according to the standard.

This is why styled display names revert on some platforms. Nothing rejected them; something normalised them. Platforms that apply NFKC to usernames and display names do it on purpose, usually to stop homograph impersonation โ€” ๐—”๐—ฑ๐—บ๐—ถ๐—ป and Admin fold to the same string, so one account cannot masquerade as the other.

It cuts the other way too. If you are storing user-generated text and you want styled input preserved, do not normalise with a compatibility form. NFC is safe for this: it leaves the mathematical alphabets alone, because they have no canonical decomposition, only a compatibility one.

Which form to use

TaskForm
Storing user textNFC
Comparing strings for equalityNFC on both sides
Search and matchingNFKC, plus case folding
Usernames, identifiers, domainsNFKC โ€” deliberately lossy, to stop spoofing
Stripping accentsNFD, then remove U+0300โ€“U+036F
Preserving styled textNFC, never NFKC
File namesNFC, and be aware macOS differs

Normalise for comparison, store what the user typed

That table is the short answer. Keep the original NFC string in the database and build a separate NFKC-folded column for searching. That way a search for โ€œHelloโ€ finds ๐—›๐—ฒ๐—น๐—น๐—ผ and the display still shows what the user wrote.

Unicode normalization in code

Most languages have this built in.

// JavaScript
'e\u0301'.normalize('NFC') === '\u00E9';   // true
'๐—›๐—ฒ๐—น๐—น๐—ผ'.normalize('NFKC');                  // "Hello"

// strip diacritics
'caf\u00E9'.normalize('NFD').replace(/[\u0300-\u036F]/g, '');  // "cafe"

JavaScript, Python, Java and PostgreSQL

Python has unicodedata.normalize('NFC', s). Java has java.text.Normalizer. Rust needs the unicode-normalization crate. PostgreSQL has a normalize() function since version 13.

Two cautions

Normalisation is not idempotent across forms โ€” normalising to NFC then NFKC is not the same as going straight to NFKC in every case, so pick one and apply it once. And normalisation is not case folding; if you want case-insensitive comparison you need both, in that order.

What normalisation does not fix

Zalgo. NFC combines a base and a mark only where a precomposed character exists. There is no precomposed form of e with twenty-nine marks, so NFC tidies the first one and leaves the rest. Normalisation is not a Zalgo filter โ€” see what Zalgo text is.

Invisible characters. Zero-width spaces and joiners survive every normalisation form. They have to: they are meaningful in scripts that need them. Removing them is a separate filtering step, covered in the zero-width space reference.

Homographs across scripts. Cyrillic ะฐ (U+0430) and Latin a (U+0061) look identical and are not equivalent in any form, because they genuinely are different letters. Defending against that needs Unicode Technical Standard #39, not normalisation.

Quick answers

Which normalisation form should I use by default?
NFC, for storage and for comparison.

Is NFKC reversible?
No. It discards information permanently.

Why did my styled username revert to plain text?
Something in the pipeline applied NFKC. That is the standard behaviour for those characters.

Does normalisation change the length of a string?
Often yes. NFD usually lengthens it, NFC usually shortens it, NFKC can do either.

Do I need to normalise if all my text is ASCII?
No. ASCII is unchanged by every form.