Unicode and text test values
Characters that look innocent and break counting, comparison and display. 12 values, each with what it breaks and what a correct application does instead.
Type them one at a time with AltShiftN
in the palette, or take them all into a script with nkb emit unicode-text.
Right-to-left override
unicode-text/rtl-override
raporttxt.exe
raport\u202Etxt.exe
- 1 character that changes the direction of text
- 14 graphemes
- 14 code points
- 16 bytes
- 14 UTF-16 units
- What it breaks
- A direction-override character makes the text render reversed: the value shown to the user is not the value stored.
- What a correct application does
- Stripped, escaped, or displayed with the override neutralised - never rendered as-is.
Soft hyphen
unicode-text/soft-hyphen
Kowalski
Kowal\u00ADski
- 1 soft hyphen
- 9 graphemes
- 9 code points
- 10 bytes
- 9 UTF-16 units
- What it breaks
- An invisible hyphen inside a word defeats exact search and makes two identical-looking names different records.
- What a correct application does
- Stored unchanged and found by a search for Kowalski, or normalised away on save.
Combining acute accent
unicode-text/combining-acute
é
é
- 1 grapheme
- 2 code points
- 3 bytes
- 2 UTF-16 units
- What it breaks
- The same letter written as a base plus a combining mark compares as different from the single-character form.
- What a correct application does
- Normalised on save, or compared in a normalisation-aware way.
Stacked combining marks
unicode-text/zalgo
è́̂̃̄̅̆̇̈̉
è́̂̃̄̅̆̇̈̉
- 1 grapheme
- 11 code points
- 21 bytes
- 11 UTF-16 units
- What it breaks
- Stacked combining marks overflow the line vertically and cover the interface around the field.
- What a correct application does
- Accepted and clipped to the box of the field, or rejected with a limit on combining marks.
Byte-order mark inside a value
unicode-text/bom-inside
JanKowalski
Jan\uFEFFKowalski
- 1 zero-width character
- 12 graphemes
- 12 code points
- 14 bytes
- 12 UTF-16 units
- What it breaks
- A byte-order mark in the middle of a value survives every copy and silently breaks exact comparison.
- What a correct application does
- Stripped on save, or preserved and made visible in the interface.
Bell control character
unicode-text/bell-control
JanKowalski
Jan\u0007Kowalski
- 1 other control character
- 12 graphemes
- 12 code points
- 12 bytes
- 12 UTF-16 units
- What it breaks
- A control character passed to a terminal or a log makes the output do something instead of saying something.
- What a correct application does
- Rejected or escaped before it reaches any log, terminal or report.
Full-width Latin letters
unicode-text/fullwidth-latin
fullwidth
fullwidth
- 9 graphemes
- 9 code points
- 27 bytes
- 9 UTF-16 units
- What it breaks
- Full-width letters read as Latin to a human and as different characters to search, filters and uniqueness checks.
- What a correct application does
- Normalised for comparison, or accepted and consistently treated as distinct.
Mathematical bold letters
unicode-text/math-bold
𝐇𝐞𝐥𝐥𝐨
𝐇𝐞𝐥𝐥𝐨
- 5 graphemes
- 5 code points
- 20 bytes
- 10 UTF-16 units
- What it breaks
- Mathematical letter variants defeat profanity filters and search while looking like ordinary words.
- What a correct application does
- Normalised for comparison and filtering, or rejected in fields where plain text is expected.
Upside-down letters
unicode-text/upside-down
ʇxǝʇ
ʇxǝʇ
- 4 graphemes
- 4 code points
- 7 bytes
- 4 UTF-16 units
- What it breaks
- Ordinary letters from another block, used to bypass filters and to break line rendering in lists.
- What a correct application does
- Accepted as plain text without breaking the layout, and covered by the same filters as normal text.
Mixed CJK characters
unicode-text/cjk-mixed
名前テスト
名前テスト
- 5 graphemes
- 5 code points
- 15 bytes
- 5 UTF-16 units
- What it breaks
- Double-width characters break column alignment and confuse character counters that assume one character is one cell.
- What a correct application does
- Rendered and counted correctly, with column widths adapting instead of overflowing.
Arabic inside a Latin sentence
unicode-text/arabic-in-latin
JanمرحباKowalski
Jan مرحبا Kowalski
- 18 graphemes
- 18 code points
- 23 bytes
- 18 UTF-16 units
- What it breaks
- Right-to-left text inside a left-to-right sentence reorders on screen, so the visual order is not the stored order.
- What a correct application does
- Stored in logical order and rendered with correct isolation, so copying returns what was typed.
Emoji in a name field
unicode-text/emoji-in-name
Jan😀Kowalski
Jan😀Kowalski
- 12 graphemes
- 12 code points
- 15 bytes
- 13 UTF-16 units
- What it breaks
- A character outside the basic plane in a name field breaks byte-length limits and databases configured for narrow encodings.
- What a correct application does
- Stored intact or rejected with a message - never truncated mid-character.