What a Unicode Code Point Is, and Why It Is Not a Byte

A Unicode code point explained: what U+0041 means, how code points differ from bytes and characters, and why a five-letter word can be twenty characters long.

Concepts7 min read
Two code points beside their characters: U plus 0041 equals A, U plus 1D400 equals a bold A

A code point is a number that Unicode has assigned to one character. That is the entire definition, and almost every confusion about text comes from mixing it up with the two things either side of it: the bytes underneath, and the glyph on screen.

The notation

A code point is written U+ followed by at least four hexadecimal digits:

Code pointCharacterName
U+0041ALATIN CAPITAL LETTER A
U+0061aLATIN SMALL LETTER A
U+00E9รฉLATIN SMALL LETTER E WITH ACUTE
U+0301โ—ŒฬCOMBINING ACUTE ACCENT
U+1D400๐€MATHEMATICAL BOLD CAPITAL A
U+16A0แš RUNIC LETTER FEHU FEOH FE F
U+2800โ €BRAILLE PATTERN BLANK

The hex digits are the number. U+0041 is 65 in decimal, which is why A is 65 in ASCII too โ€” Unicodeโ€™s first 128 code points were deliberately made identical to ASCII so existing text would keep working.

Unicode character names are permanent

Names are in capitals by convention, and they are immutable. Once Unicode assigns a name it never changes it, even when the name turns out to be wrong. U+FE18 is officially PRESENTATION FORM FOR VERTICAL RIGHT WHITE LENTICULAR BRAKCET โ€” the typo is in the standard forever, with a formal correction note beside it.

The codespace

Unicodeโ€™s range runs from U+0000 to U+10FFFF, which is 1,114,112 possible code points. Roughly 150,000 are currently assigned; the rest are unassigned, reserved, or set aside for private use.

The space is divided into 17 planes of 65,536 each:

PlaneRangeWhat lives there
0 โ€” Basic Multilingual PlaneU+0000โ€“U+FFFFAlmost every modern script
1 โ€” Supplementary MultilingualU+10000โ€“U+1FFFFHistoric scripts, emoji, the mathematical alphabets
2 โ€” Supplementary IdeographicU+20000โ€“U+2FFFFRare CJK ideographs
3 โ€” Tertiary IdeographicU+30000โ€“U+3FFFFMore CJK
14 โ€” Supplementary Special-purposeU+E0000โ€“U+E01FFTags, variation selectors
15โ€“16 โ€” Private UseU+F0000โ€“U+10FFFDWhatever an organisation wants

That split has practical consequences. Plane 0 code points fit in 16 bits; everything above needs more, which is where surrogate pairs come from โ€” and where a lot of software still has bugs.

Code point, character, glyph

Three different things, routinely confused.

Code point โ€” a number Unicode assigned. U+0065 and U+0301.

Character โ€” what a user thinks of as one thing. รฉ is one character to a reader, but it can be one code point (U+00E9) or two (U+0065 U+0301). Both are correct and both are รฉ.

Glyph โ€” the shape a font draws. One code point can produce several glyphs depending on context, and several code points can produce one glyph. The Arabic letter that changes shape at the start, middle and end of a word is one code point and four glyphs. A ligature like fi is two code points and one glyph.

Unicode has a fourth term for the useful unit in the middle: a grapheme cluster, defined in Unicode Standard Annex #29, is the sequence of code points that a reader would call one character. It is what a text cursor should move over, what a backspace should delete, and what โ€œlengthโ€ should usually count โ€” which is why ๐Ÿ‘ฉโ€๐Ÿ‘ฉโ€๐Ÿ‘งโ€๐Ÿ‘ฆ, a family emoji built from seven code points joined by zero-width joiners, deletes in one press in software that gets this right.

Why your five-letter word is twenty characters

This is the most common practical consequence, and the character counter on this site exists mostly because of it.

Take โ€œHelloโ€ styled as bold sans: ๐—›๐—ฒ๐—น๐—น๐—ผ. That is five code points โ€” U+1D5DB, U+1D5F2, U+1D5F9, U+1D5F9, U+1D5FC โ€” which is fine. But every one of them is above U+FFFF, so in UTF-16, which is what JavaScript and most platform APIs use internally, each takes two 16-bit units. Naive length code reports 10.

Now take Zalgo text. Zฬปอ–ฤƒฬ‘อŠlฬทอ‚อ˜gฬขฬ–ฬ•oฬขฬ–ฬ• looks like five letters. It is five base code points plus every combining mark stacked on them โ€” easily forty or fifty code points, and a bio field with a 150-character limit will reject it while showing you five letters.

And a platform might count any of three things: code points, UTF-16 units, or bytes. Twitter counts weighted code points. Instagramโ€™s bio limit behaves like code points. A database column declared VARCHAR(150) may well be counting bytes. The same string can be legal in one and rejected by the next.

Reading and writing code points

Unicode code point lookup, and typing one

Finding one. The Unicode Character Database is the authoritative source, and the code charts are the browsable version. Every character on this site links to the chart for its block.

Typing one. On Windows, type the hex digits and press Alt+X in applications that support it. On macOS, the Unicode Hex Input keyboard takes Option plus the digits. On Linux, Ctrl+Shift+U then the digits. In a web page, 𝐀 in HTML and \u{1D400} in JavaScript and CSS.

Counting them correctly. In JavaScript, "๐€".length is 2 and [..."๐€"].length is 1, because spreading a string iterates code points rather than UTF-16 units. For grapheme clusters, Intl.Segmenter is the right tool.

Code point ranges used on this site

Every style here is a mapping from one range of code points to another. The Unicode blocks reference lists all of them; the largest are:

RangeBlockUsed for
U+1D400โ€“U+1D7FFMathematical Alphanumeric SymbolsBold, italic, script, fraktur, double-struck, monospace
U+2460โ€“U+24FFEnclosed AlphanumericsBubble letters
U+FF00โ€“U+FFEFHalfwidth and Fullwidth FormsVaporwave spacing
U+1D00โ€“U+1D7FPhonetic ExtensionsSmall capitals
U+0300โ€“U+036FCombining Diacritical MarksZalgo, underline, strikethrough
U+16A0โ€“U+16FFRunicThe rune translator
U+2800โ€“U+28FFBraille PatternsThe braille blank on /invisible-character

Quick answers

Is a code point the same as a byte?
No. How many bytes a code point takes depends on the encoding โ€” 1 to 4 in UTF-8, 2 or 4 in UTF-16. See UTF-8 vs Unicode.

What does U+ stand for?
Nothing expanded โ€” it is just the standard prefix marking what follows as a Unicode code point in hexadecimal.

Why are some code points unassigned?
Unicode leaves room deliberately, so related characters can be added near each other later.

Can a code point be removed?
No. Unicodeโ€™s stability policy means an assigned code point keeps its meaning forever. Characters can be deprecated but never unassigned.

What are surrogates?
Code points U+D800โ€“U+DFFF, reserved so UTF-16 can represent characters above U+FFFF as a pair. They are never valid on their own, and a lone surrogate is the source of a lot of corrupted text.