Unicode and text test values

Characters that look innocent and break counting, comparison and display. 12 values, each with what it breaks and what a correct application does instead.

unicode-text version 1.0, updated 2026-09-08 CC-BY-4.0

Type them one at a time with AltShiftN in the palette, or take them all into a script with nkb emit unicode-text.

Right-to-left override

unicode-text/rtl-override
raporttxt.exe raport\u202Etxt.exe
  • 1 character that changes the direction of text
  • 14 graphemes
  • 14 code points
  • 16 bytes
  • 14 UTF-16 units
What it breaks
A direction-override character makes the text render reversed: the value shown to the user is not the value stored.
What a correct application does
Stripped, escaped, or displayed with the override neutralised - never rendered as-is.

Soft hyphen

unicode-text/soft-hyphen
Kowalski Kowal\u00ADski
  • 1 soft hyphen
  • 9 graphemes
  • 9 code points
  • 10 bytes
  • 9 UTF-16 units
What it breaks
An invisible hyphen inside a word defeats exact search and makes two identical-looking names different records.
What a correct application does
Stored unchanged and found by a search for Kowalski, or normalised away on save.

Combining acute accent

unicode-text/combining-acute
é é
  • 1 grapheme
  • 2 code points
  • 3 bytes
  • 2 UTF-16 units
What it breaks
The same letter written as a base plus a combining mark compares as different from the single-character form.
What a correct application does
Normalised on save, or compared in a normalisation-aware way.

Stacked combining marks

unicode-text/zalgo
è́̂̃̄̅̆̇̈̉ è́̂̃̄̅̆̇̈̉
  • 1 grapheme
  • 11 code points
  • 21 bytes
  • 11 UTF-16 units
What it breaks
Stacked combining marks overflow the line vertically and cover the interface around the field.
What a correct application does
Accepted and clipped to the box of the field, or rejected with a limit on combining marks.

Byte-order mark inside a value

unicode-text/bom-inside
JanKowalski Jan\uFEFFKowalski
  • 1 zero-width character
  • 12 graphemes
  • 12 code points
  • 14 bytes
  • 12 UTF-16 units
What it breaks
A byte-order mark in the middle of a value survives every copy and silently breaks exact comparison.
What a correct application does
Stripped on save, or preserved and made visible in the interface.

Bell control character

unicode-text/bell-control
JanKowalski Jan\u0007Kowalski
  • 1 other control character
  • 12 graphemes
  • 12 code points
  • 12 bytes
  • 12 UTF-16 units
What it breaks
A control character passed to a terminal or a log makes the output do something instead of saying something.
What a correct application does
Rejected or escaped before it reaches any log, terminal or report.

Full-width Latin letters

unicode-text/fullwidth-latin
fullwidth fullwidth
  • 9 graphemes
  • 9 code points
  • 27 bytes
  • 9 UTF-16 units
What it breaks
Full-width letters read as Latin to a human and as different characters to search, filters and uniqueness checks.
What a correct application does
Normalised for comparison, or accepted and consistently treated as distinct.

Mathematical bold letters

unicode-text/math-bold
𝐇𝐞𝐥𝐥𝐨 𝐇𝐞𝐥𝐥𝐨
  • 5 graphemes
  • 5 code points
  • 20 bytes
  • 10 UTF-16 units
What it breaks
Mathematical letter variants defeat profanity filters and search while looking like ordinary words.
What a correct application does
Normalised for comparison and filtering, or rejected in fields where plain text is expected.

Upside-down letters

unicode-text/upside-down
ʇxǝʇ ʇxǝʇ
  • 4 graphemes
  • 4 code points
  • 7 bytes
  • 4 UTF-16 units
What it breaks
Ordinary letters from another block, used to bypass filters and to break line rendering in lists.
What a correct application does
Accepted as plain text without breaking the layout, and covered by the same filters as normal text.

Mixed CJK characters

unicode-text/cjk-mixed
名前テスト 名前テスト
  • 5 graphemes
  • 5 code points
  • 15 bytes
  • 5 UTF-16 units
What it breaks
Double-width characters break column alignment and confuse character counters that assume one character is one cell.
What a correct application does
Rendered and counted correctly, with column widths adapting instead of overflowing.

Arabic inside a Latin sentence

unicode-text/arabic-in-latin
JanمرحباKowalski Jan مرحبا Kowalski
  • 18 graphemes
  • 18 code points
  • 23 bytes
  • 18 UTF-16 units
What it breaks
Right-to-left text inside a left-to-right sentence reorders on screen, so the visual order is not the stored order.
What a correct application does
Stored in logical order and rendered with correct isolation, so copying returns what was typed.

Emoji in a name field

unicode-text/emoji-in-name
Jan😀Kowalski Jan😀Kowalski
  • 12 graphemes
  • 12 code points
  • 15 bytes
  • 13 UTF-16 units
What it breaks
A character outside the basic plane in a name field breaks byte-length limits and databases configured for narrow encodings.
What a correct application does
Stored intact or rejected with a message - never truncated mid-character.