Unicode Normalization: NFC, NFD, NFKC and NFKD Explained
Unicode normalization forms compared: what NFC, NFD, NFKC and NFKD each do, which one to use for storage and search, and why NFKC destroys styled text.

Unicode lets you write the same character more than one way. รฉ can be a single code point, U+00E9, or two โ e at U+0065 followed by COMBINING ACUTE ACCENT at U+0301. Both are correct, both look identical, and neither is equal to the other in a byte comparison.
Normalization is the process of converting text to one canonical form so that comparisons work. There are four forms, and choosing the wrong one silently destroys data.
Why multiple representations exist
Historical compatibility, mostly. When Unicode absorbed the existing national character sets of the 1980s and 1990s, it took their precomposed characters with them so round-tripping would be lossless. Latin-1 had a single character for รฉ, so Unicode has one. But Unicode also has the general mechanism of base plus combining mark, which can produce the same result.
So there are now two encodings of รฉ, and there will be forever โ Unicodeโs stability policy forbids removing either.
The general case is worse than two. Vietnamese แบฟ can be written four ways. Korean Hangul syllables can be one precomposed character or a sequence of conjoining jamo. Any text arriving from an unknown source may be in any of them.
The two axes
The four forms are two decisions crossed with each other.
Composition. Do you want the precomposed single character (composed, C) or the base-plus-marks sequence (decomposed, D)?
Equivalence. Do you apply only canonical equivalence โ sequences that genuinely are the same character โ or also compatibility equivalence (K), which folds characters that mean the same thing but look different?
| Canonical | Compatibility | |
|---|---|---|
| Composed | NFC | NFKC |
| Decomposed | NFD | NFKD |
The rules are defined in Unicode Standard Annex #15, which is the authoritative document.
Canonical equivalence: NFC and NFD
Canonical equivalence means the sequences are the same character, differently written. Converting between them is lossless in both directions.
| Input | NFC | NFD |
|---|---|---|
รฉ (U+00E9) | U+00E9 | U+0065 U+0301 |
e + U+0301 | U+00E9 | U+0065 U+0301 |
ร
(U+00C5) | U+00C5 | U+0041 U+030A |
ร
(U+212B, angstrom sign) | U+00C5 | U+0041 U+030A |
That last row is worth a look. U+212B ANGSTROM SIGN is canonically equivalent to ร
โ the standard says they are the same character โ so normalising in either form collapses it. This is the correct behaviour, and it surprises people.
NFC is the default choice. It is what the W3C recommends for the web, what most text arrives in already, and what you want in a database. NFD is mainly useful when you need to strip diacritics: decompose, then drop everything in the combining marks range, and cafรฉ becomes cafe.
macOS is the notable oddity โ its file system stores names in a variant of NFD, which is why a filename copied from a Mac can compare unequal to the same filename typed on Linux.
Compatibility equivalence: NFKC and NFKD
Compatibility equivalence is much more aggressive. It folds characters that are semantically the same but visually distinct, and it is lossy. You cannot undo it.
| Input | After NFKC | What was lost |
|---|---|---|
๏ฌ (U+FB01 ligature) | fi | the ligature |
ยฒ (superscript two) | 2 | the superscript |
โ
ฃ (Roman numeral four) | IV | the single-character numeral |
ใ (square kg) | kg | the compact form |
๏ผจ๏ฝ
๏ฝ๏ฝ๏ฝ (fullwidth) | Hello | the fullwidth forms |
๐๐ฒ๐น๐น๐ผ (bold sans) | Hello | all the styling |
แดดแตหกหกแต (superscript) | Hello | the superscript forms |
โฝโโโโ (circled) | (H)(e)(l)(l)(o) | the circles |
That table is the single most important thing on this page for anyone using styled text.
Why NFKC destroys styled text
Every style on this site is a mapping into a Unicode block whose characters carry a compatibility decomposition back to plain ASCII. That is not a flaw in the styles; it is what the standard says those characters are. U+1D5DB MATHEMATICAL SANS-SERIF BOLD CAPITAL H is formally defined as a compatibility variant of H, because in a formula it is an H, printed differently.
So any system that applies NFKC to incoming text will turn ๐๐ฒ๐น๐น๐ผ into Hello and ๐ฎ๐ถ๐๐ถ๐ฝ into Sarah. Not corrupt it โ convert it, correctly and deliberately, according to the standard.
This is why styled display names revert on some platforms. Nothing rejected them; something normalised them. Platforms that apply NFKC to usernames and display names do it on purpose, usually to stop homograph impersonation โ ๐๐ฑ๐บ๐ถ๐ป and Admin fold to the same string, so one account cannot masquerade as the other.
It cuts the other way too. If you are storing user-generated text and you want styled input preserved, do not normalise with a compatibility form. NFC is safe for this: it leaves the mathematical alphabets alone, because they have no canonical decomposition, only a compatibility one.
Which form to use
| Task | Form |
|---|---|
| Storing user text | NFC |
| Comparing strings for equality | NFC on both sides |
| Search and matching | NFKC, plus case folding |
| Usernames, identifiers, domains | NFKC โ deliberately lossy, to stop spoofing |
| Stripping accents | NFD, then remove U+0300โU+036F |
| Preserving styled text | NFC, never NFKC |
| File names | NFC, and be aware macOS differs |
Normalise for comparison, store what the user typed
That table is the short answer. Keep the original NFC string in the database and build a separate NFKC-folded column for searching. That way a search for โHelloโ finds ๐๐ฒ๐น๐น๐ผ and the display still shows what the user wrote.
Unicode normalization in code
Most languages have this built in.
// JavaScript
'e\u0301'.normalize('NFC') === '\u00E9'; // true
'๐๐ฒ๐น๐น๐ผ'.normalize('NFKC'); // "Hello"
// strip diacritics
'caf\u00E9'.normalize('NFD').replace(/[\u0300-\u036F]/g, ''); // "cafe"
JavaScript, Python, Java and PostgreSQL
Python has unicodedata.normalize('NFC', s). Java has java.text.Normalizer. Rust needs the unicode-normalization crate. PostgreSQL has a normalize() function since version 13.
Two cautions
Normalisation is not idempotent across forms โ normalising to NFC then NFKC is not the same as going straight to NFKC in every case, so pick one and apply it once. And normalisation is not case folding; if you want case-insensitive comparison you need both, in that order.
What normalisation does not fix
Zalgo. NFC combines a base and a mark only where a precomposed character exists. There is no precomposed form of e with twenty-nine marks, so NFC tidies the first one and leaves the rest. Normalisation is not a Zalgo filter โ see what Zalgo text is.
Invisible characters. Zero-width spaces and joiners survive every normalisation form. They have to: they are meaningful in scripts that need them. Removing them is a separate filtering step, covered in the zero-width space reference.
Homographs across scripts. Cyrillic ะฐ (U+0430) and Latin a (U+0061) look identical and are not equivalent in any form, because they genuinely are different letters. Defending against that needs Unicode Technical Standard #39, not normalisation.
Quick answers
Which normalisation form should I use by default?
NFC, for storage and for comparison.
Is NFKC reversible?
No. It discards information permanently.
Why did my styled username revert to plain text?
Something in the pipeline applied NFKC. That is the standard behaviour for those characters.
Does normalisation change the length of a string?
Often yes. NFD usually lengthens it, NFC usually shortens it, NFKC can do either.
Do I need to normalise if all my text is ASCII?
No. ASCII is unchanged by every form.
Related
- What a Unicode code point is โ the units normalisation operates on
- UTF-8 vs Unicode โ encoding, which is a separate concern
- Combining diacritical marks โ the marks NFD produces
- Mathematical alphanumeric symbols โ the block NFKC folds away