What a Unicode Code Point Is, and Why It Is Not a Byte
A Unicode code point explained: what U+0041 means, how code points differ from bytes and characters, and why a five-letter word can be twenty characters long.

A code point is a number that Unicode has assigned to one character. That is the entire definition, and almost every confusion about text comes from mixing it up with the two things either side of it: the bytes underneath, and the glyph on screen.
The notation
A code point is written U+ followed by at least four hexadecimal digits:
| Code point | Character | Name |
|---|---|---|
| U+0041 | A | LATIN CAPITAL LETTER A |
| U+0061 | a | LATIN SMALL LETTER A |
| U+00E9 | รฉ | LATIN SMALL LETTER E WITH ACUTE |
| U+0301 | โฬ | COMBINING ACUTE ACCENT |
| U+1D400 | ๐ | MATHEMATICAL BOLD CAPITAL A |
| U+16A0 | แ | RUNIC LETTER FEHU FEOH FE F |
| U+2800 | โ | BRAILLE PATTERN BLANK |
The hex digits are the number. U+0041 is 65 in decimal, which is why A is 65 in ASCII too โ Unicodeโs first 128 code points were deliberately made identical to ASCII so existing text would keep working.
Unicode character names are permanent
Names are in capitals by convention, and they are immutable. Once Unicode assigns a name it never changes it, even when the name turns out to be wrong. U+FE18 is officially PRESENTATION FORM FOR VERTICAL RIGHT WHITE LENTICULAR BRAKCET โ the typo is in the standard forever, with a formal correction note beside it.
The codespace
Unicodeโs range runs from U+0000 to U+10FFFF, which is 1,114,112 possible code points. Roughly 150,000 are currently assigned; the rest are unassigned, reserved, or set aside for private use.
The space is divided into 17 planes of 65,536 each:
| Plane | Range | What lives there |
|---|---|---|
| 0 โ Basic Multilingual Plane | U+0000โU+FFFF | Almost every modern script |
| 1 โ Supplementary Multilingual | U+10000โU+1FFFF | Historic scripts, emoji, the mathematical alphabets |
| 2 โ Supplementary Ideographic | U+20000โU+2FFFF | Rare CJK ideographs |
| 3 โ Tertiary Ideographic | U+30000โU+3FFFF | More CJK |
| 14 โ Supplementary Special-purpose | U+E0000โU+E01FF | Tags, variation selectors |
| 15โ16 โ Private Use | U+F0000โU+10FFFD | Whatever an organisation wants |
That split has practical consequences. Plane 0 code points fit in 16 bits; everything above needs more, which is where surrogate pairs come from โ and where a lot of software still has bugs.
Code point, character, glyph
Three different things, routinely confused.
Code point โ a number Unicode assigned. U+0065 and U+0301.
Character โ what a user thinks of as one thing. รฉ is one character to a reader, but it can be one code point (U+00E9) or two (U+0065 U+0301). Both are correct and both are รฉ.
Glyph โ the shape a font draws. One code point can produce several glyphs depending on context, and several code points can produce one glyph. The Arabic letter that changes shape at the start, middle and end of a word is one code point and four glyphs. A ligature like fi is two code points and one glyph.
Unicode has a fourth term for the useful unit in the middle: a grapheme cluster, defined in Unicode Standard Annex #29, is the sequence of code points that a reader would call one character. It is what a text cursor should move over, what a backspace should delete, and what โlengthโ should usually count โ which is why ๐ฉโ๐ฉโ๐งโ๐ฆ, a family emoji built from seven code points joined by zero-width joiners, deletes in one press in software that gets this right.
Why your five-letter word is twenty characters
This is the most common practical consequence, and the character counter on this site exists mostly because of it.
Take โHelloโ styled as bold sans: ๐๐ฒ๐น๐น๐ผ. That is five code points โ U+1D5DB, U+1D5F2, U+1D5F9, U+1D5F9, U+1D5FC โ which is fine. But every one of them is above U+FFFF, so in UTF-16, which is what JavaScript and most platform APIs use internally, each takes two 16-bit units. Naive length code reports 10.
Now take Zalgo text. Zฬปอฤฬอlฬทออgฬขฬฬoฬขฬฬ looks like five letters. It is five base code points plus every combining mark stacked on them โ easily forty or fifty code points, and a bio field with a 150-character limit will reject it while showing you five letters.
And a platform might count any of three things: code points, UTF-16 units, or bytes. Twitter counts weighted code points. Instagramโs bio limit behaves like code points. A database column declared VARCHAR(150) may well be counting bytes. The same string can be legal in one and rejected by the next.
Reading and writing code points
Unicode code point lookup, and typing one
Finding one. The Unicode Character Database is the authoritative source, and the code charts are the browsable version. Every character on this site links to the chart for its block.
Typing one. On Windows, type the hex digits and press Alt+X in applications that support it. On macOS, the Unicode Hex Input keyboard takes Option plus the digits. On Linux, Ctrl+Shift+U then the digits. In a web page, 𝐀 in HTML and \u{1D400} in JavaScript and CSS.
Counting them correctly. In JavaScript, "๐".length is 2 and [..."๐"].length is 1, because spreading a string iterates code points rather than UTF-16 units. For grapheme clusters, Intl.Segmenter is the right tool.
Code point ranges used on this site
Every style here is a mapping from one range of code points to another. The Unicode blocks reference lists all of them; the largest are:
| Range | Block | Used for |
|---|---|---|
| U+1D400โU+1D7FF | Mathematical Alphanumeric Symbols | Bold, italic, script, fraktur, double-struck, monospace |
| U+2460โU+24FF | Enclosed Alphanumerics | Bubble letters |
| U+FF00โU+FFEF | Halfwidth and Fullwidth Forms | Vaporwave spacing |
| U+1D00โU+1D7F | Phonetic Extensions | Small capitals |
| U+0300โU+036F | Combining Diacritical Marks | Zalgo, underline, strikethrough |
| U+16A0โU+16FF | Runic | The rune translator |
| U+2800โU+28FF | Braille Patterns | The braille blank on /invisible-character |
Quick answers
Is a code point the same as a byte?
No. How many bytes a code point takes depends on the encoding โ 1 to 4 in UTF-8, 2 or 4 in UTF-16. See UTF-8 vs Unicode.
What does U+ stand for?
Nothing expanded โ it is just the standard prefix marking what follows as a Unicode code point in hexadecimal.
Why are some code points unassigned?
Unicode leaves room deliberately, so related characters can be added near each other later.
Can a code point be removed?
No. Unicodeโs stability policy means an assigned code point keeps its meaning forever. Characters can be deprecated but never unassigned.
What are surrogates?
Code points U+D800โU+DFFF, reserved so UTF-16 can represent characters above U+FFFF as a pair. They are never valid on their own, and a lone surrogate is the source of a lot of corrupted text.