Mathematical Alphanumeric Symbols: The Block Behind It All

The Mathematical Alphanumeric Symbols block, U+1D400 to U+1D7FF: every styled alphabet it contains, the sub-ranges, and the holes where letters are missing.

Character blocks8 min read
The capital A in seven forms: bold, italic, script, fraktur, double-struck, monospace and sans bold

One Unicode block does most of the work on this site. Mathematical Alphanumeric Symbols, U+1D400 to U+1D7FF, holds thirteen complete styled alphabets plus five sets of styled digits โ€” bold, italic, script, fraktur, double-struck, monospace and the sans-serif variants of each.

It was not designed for social media. Understanding why explains both what it can do and every gap in it.

Why the block exists

Mathematical notation uses typeface as meaning. In a physics paper, ๐ is a magnetic field vector, ๐ต is its magnitude, ๐”… might be a basis and ๐”น a Boolean domain. The same Latin letter B, printed four ways, means four different things โ€” and the distinction is semantic, not decorative.

Plain text could not carry that. So when Unicode set out to support mathematical publishing, it had to give each styled form its own code point, and the result was 996 characters covering every combination mathematics conventionally uses. The rationale is set out in Unicode Technical Report #25, the standardโ€™s document on mathematics.

The unintended consequence is the entire copy-and-paste font industry. These are real characters, so they survive a paste into any field that accepts text โ€” including every bio, username and comment box on the internet, none of which have a bold button.

The sub-ranges

StartAlphabetExample
U+1D400Bold (serif)๐€๐š
U+1D434Italic (serif)๐ด๐‘Ž
U+1D468Bold italic (serif)๐‘จ๐’‚
U+1D49CScript๐’œ๐’ถ
U+1D4D0Bold script๐“๐“ช
U+1D504Fraktur๐”„๐”ž
U+1D538Double-struck๐”ธ๐•’
U+1D56CBold fraktur๐•ฌ๐–†
U+1D5A0Sans-serif๐– ๐–บ
U+1D5D4Sans-serif bold๐—”๐—ฎ
U+1D608Sans-serif italic๐˜ˆ๐˜ข
U+1D63CSans-serif bold italic๐˜ผ๐™–
U+1D670Monospace๐™ฐ๐šŠ
U+1D6A4Dotless i and j๐šค๐šฅ
U+1D6A8Bold Greek๐šจ๐›‚
U+1D7CEBold digits๐ŸŽ๐Ÿ—
U+1D7D8Double-struck digits๐Ÿ˜๐Ÿก
U+1D7E2Sans-serif digits๐Ÿข๐Ÿซ
U+1D7ECSans-serif bold digits๐Ÿฌ๐Ÿต
U+1D7F6Monospace digits๐Ÿถ๐Ÿฟ

Each Latin alphabet occupies 52 consecutive code points: 26 capitals then 26 lowercase, in order. That regularity is why converting text is arithmetic โ€” subtract A, add U+1D400 โ€” and why every implementation of this produces identical output.

The holes, and where the missing letters live

Now the part that catches every implementation once.

Several letters are not in this block, because Unicode had already encoded them years earlier in Letterlike Symbols at U+2100โ€“U+214F. Unicode does not encode the same character twice, so their positions in U+1D400 are reserved and permanently unassigned.

AlphabetMissing from the blockActual code point
Script capitalsB, E, F, H, I, L, M, Rโ„ฌ โ„ฐ โ„ฑ โ„‹ โ„ โ„’ โ„ณ โ„›
Script lowercasee, g, oโ„ฏ โ„Š โ„ด
Fraktur capitalsC, H, I, R, Zโ„ญ โ„Œ โ„‘ โ„œ โ„จ
Double-struck capitalsC, H, N, P, Q, R, Zโ„‚ โ„ โ„• โ„™ โ„š โ„ โ„ค
Italic lowercasehโ„Ž

You can see the effect in real output. Our script style renders โ€œHelloโ€ as โ„‹โ„ฏ๐“๐“โ„ด โ€” the H and the o come from Letterlike Symbols at U+210B and U+2134, the lโ€™s from U+1D4C1 in the mathematical block. Two different blocks, one word, and it only looks right because the converter knows about the exceptions.

A converter that does naive arithmetic produces an unassigned code point for those letters instead, which renders as an empty box. If you have ever seen a script-style generator where every H is a rectangle, this is why.

What the block does not contain

No cursive or fraktur numbers

No styled digits for most alphabets. There are bold, double-struck, sans, sans-bold and monospace digits, and that is the complete list. Script, fraktur, italic and bold-italic have no digits at all, so any converter passes them through unchanged. A phone number in โ€œcursiveโ€ is a plain phone number.

No styled punctuation. Not one mark. Commas, full stops, apostrophes, brackets โ€” all pass through as ordinary ASCII, in every style.

No accented or non-Latin letters

No accented letters. No ๐žฬ, no ๐งฬƒ, nothing. Any language that uses diacritics falls back to plain characters the moment it needs one, which is why styled text in French, Spanish, Polish or Vietnamese always looks half-converted.

No lowercase in some neighbouring styles. This block is complete on that front, but styles built elsewhere โ€” squared, inverted circled โ€” are not. See the Unicode blocks reference.

Compatibility decomposition: how it can be undone

Every character in this block carries a compatibility decomposition to its plain ASCII equivalent. U+1D5D4 MATHEMATICAL SANS-SERIF BOLD CAPITAL A formally decomposes to A, because in mathematical terms it is an A, printed differently.

That has a direct consequence for anyone using these characters: NFKC normalisation converts them all back to plain letters. ๐—›๐—ฒ๐—น๐—น๐—ผ becomes Hello, exactly as the standard specifies.

Platforms apply NFKC to usernames and identifiers on purpose, to stop one account impersonating another with lookalike characters. If your styled display name keeps reverting to plain text, nothing rejected it โ€” something normalised it. The normalisation reference has the detail.

The same property is what makes these characters searchable in a well-built system: normalise the query and the index with NFKC and a search for โ€œHelloโ€ finds ๐—›๐—ฒ๐—น๐—น๐—ผ. Most systems do not.

Font coverage

These are plane 1 characters, above U+FFFF, which used to mean patchy support. It no longer does, mostly:

  • Windows โ€” Segoe UI Symbol and Cambria Math cover the block; good since Windows 7.
  • macOS and iOS โ€” covered by the system fonts; reliable for many versions.
  • Android โ€” Noto Sans Math covers it; good on recent releases, patchy below Android 8.
  • Linux โ€” depends on whether the Noto or DejaVu math fonts are installed.
  • Smart TVs, e-readers, car displays, older smartwatches โ€” frequently not covered.

Where a font is missing the character, you get an empty rectangle, and there is nothing the sender can do about it. Why Unicode fonts break covers the fallback chain.

Accessibility

The Unicode Consortiumโ€™s own FAQ states directly that these characters are intended for mathematical notation and are not a substitute for rich text formatting.

The practical reason is screen readers. Announcing ๐—›๐—ฒ๐—น๐—น๐—ผ correctly would require the reader to recognise the compatibility decomposition; most read the character names instead, so a user hears โ€œmathematical sans-serif bold capital H, mathematical sans-serif bold small eโ€ one character at a time, or nothing at all.

So the rule that keeps this usable: style a name or a heading, keep everything that carries information in plain letters.

Styles on this site that use this block

Bold, italic, cursive script, old English / fraktur, serif, monospace and the double-struck outline style all map into the ranges above. The full font generator shows every one applied to your own text at once.

Quick answers

How many characters are in the block?
996 assigned out of 1,024 positions. The gaps are the letters encoded in Letterlike Symbols.

Why is my H a box in script style?
Because a converter used arithmetic instead of the exception table, and landed on an unassigned code point. U+210B is the correct script H.

Can I get styled numbers in cursive?
No. Script has no digits in Unicode.

Why do accented letters not convert?
The block has no accented characters. There is no styled รฉ to convert to.

Is this block the same as emoji?
No, though both live in plane 1. Emoji are in separate blocks with their own rules.