UTF-8 vs Unicode: The Difference, and Why It Matters

UTF-8 vs Unicode explained: Unicode assigns numbers to characters, UTF-8 turns those numbers into bytes, and mixing them up is what produces mojibake.

Concepts7 min read
Three characters beside their UTF-8 bytes: A is one byte, e-acute two, a bold A four

“UTF-8 vs Unicode” is a comparison between two things that are not alternatives. Unicode is a catalogue of characters with a number for each. UTF-8 is one way of writing those numbers as bytes. You need both, and they do different jobs.

The short version: Unicode says A is number 65. UTF-8 says number 65 is stored as the single byte 0x41.

What Unicode does

Unicode assigns every character a code point — a number written as U+ plus hexadecimal, such as U+0041 for A. It also gives each one a permanent name, a set of properties (is it a letter? a digit? which direction does it run?), and rules for how it behaves in text.

What Unicode deliberately does not specify is how those numbers are stored. That is an encoding’s job, and the standard defines three of them.

The three Unicode encodings

EncodingBytes per characterNotes
UTF-81 to 4ASCII-compatible; the web standard
UTF-162 or 4Used internally by JavaScript, Java, Windows, .NET
UTF-32Always 4Simple, wasteful, rarely used for storage

All three represent exactly the same set of characters. A string encoded in any of them can be converted to either of the others without loss. They differ in size, in whether they are ASCII-compatible, and in how painful they are to index into.

How UTF-8 works

UTF-8 is variable-length, and the design is the reason it won. The encoding is specified in Chapter 3 of the Unicode Standard and, for the internet, in RFC 3629.

Code point rangeBytesByte pattern
U+0000–U+007F10xxxxxxx
U+0080–U+07FF2110xxxxx 10xxxxxx
U+0800–U+FFFF31110xxxx 10xxxxxx 10xxxxxx
U+10000–U+10FFFF411110xxx 10xxxxxx 10xxxxxx 10xxxxxx

Three properties fall out of that table, and each one is why UTF-8 is everywhere:

It is backwards-compatible with ASCII. Any file containing only ASCII is already valid UTF-8, byte for byte. Fifty years of existing text needed no conversion.

It is self-synchronising. Continuation bytes always start with 10, which no leading byte does. Drop into the middle of a UTF-8 stream and you can find the next character boundary by scanning forward a byte or two. UTF-16 cannot do this.

It has no byte order. Bytes come in a defined order, so there is no big-endian and little-endian variant, and no byte order mark is needed. A UTF-8 BOM exists but is discouraged, and it causes real trouble in shell scripts and CSV files.

What it costs

Characters in the Latin alphabet cost 1 byte. Accented Latin, Greek and Cyrillic cost 2. Most CJK characters cost 3, so Chinese and Japanese text is measurably larger in UTF-8 than in UTF-16. Emoji and the mathematical alphabets cost 4.

That last row matters for anything on this site. A bold sans 𝗔 is U+1D5D4, above U+FFFF, so it is four bytes in UTF-8 against one for a plain A. A styled name that looks the same length as a plain one can be four times the size — which is exactly how you blow past a database column limit or a byte-counted API field. The character counter shows characters, code points and bytes side by side for this reason.

Mojibake: what goes wrong

Mojibake — Japanese for “character transformation” — is text encoded one way and decoded another. The classic case:

The character é is U+00E9. In UTF-8 that is two bytes, C3 A9. Decode those two bytes as Windows-1252, a single-byte encoding, and you get two characters: à and ©. So café becomes café.

Once you have seen the pattern, you can read it. é is a UTF-8 é read as Latin-1. ’ is a UTF-8 curly apostrophe read the same way.  at the start of a file is a UTF-8 BOM read as Latin-1.

Both patterns are catalogued in Wikipedia’s article on mojibake, which is a faster way to identify one than reasoning it out.

The other failure is the replacement character, � (U+FFFD), which appears when a decoder meets bytes that are not valid in the encoding it was told to use. Unlike mojibake, this one is unrecoverable: the original bytes were thrown away and replaced.

And the third is tofu, □, an empty rectangle. This one is not an encoding problem at all — the decoding worked perfectly, and the font simply has no glyph for the character. Why Unicode fonts break goes into that case.

Which to use, in practice

For files, web pages, APIs, databases, everything you store or send: UTF-8. It is the default on the web — the HTML standard requires documents to be UTF-8 — the default on Linux and macOS, and the default in every modern protocol. Declare it explicitly:

<meta charset="utf-8">
Content-Type: text/plain; charset=utf-8

In memory, you often do not get a choice. JavaScript strings are UTF-16. So are Java’s, C#‘s, and Windows’ native APIs. This is history, not preference: those platforms adopted 16-bit characters when Unicode still fit in 16 bits, and were stuck when it grew past U+FFFF in 1996.

That history is why "𝐀".length is 2 in JavaScript. The string holds one character, two UTF-16 units, and .length counts units. Iterating with [..."𝐀"] or for...of gives you code points instead, and Intl.Segmenter gives you what a reader would call characters.

One MySQL warning, because it still catches people: MySQL’s utf8 character set is not UTF-8. It is a three-byte subset that cannot store anything above U+FFFF, which means no emoji and none of the mathematical alphabets. The real thing is called utf8mb4, and it is what any column holding user text should use.

Quick answers

Is UTF-8 part of Unicode?
Yes. Unicode defines the characters and the encodings; UTF-8 is one of the encodings it specifies.

Which is better, UTF-8 or UTF-16?
UTF-8 for anything stored or transmitted. UTF-16 is not chosen so much as inherited from platform APIs.

How many bytes is an emoji?
Four in UTF-8 for a single code point, and multiples of that for sequences — a family emoji built from several people joined by zero-width joiners can run past twenty bytes.

Does UTF-8 need a BOM?
No, and adding one causes more problems than it solves.

What is UTF-7?
An obsolete seven-bit encoding for old mail systems. It is a security hazard and should not be used.