Normalize confusable characters to safe text
Normalizing confusables replaces every look-alike character with the ordinary character it imitates, so two strings that look the same also compare as equal. Do it before comparing usernames, matching blocklists, or deduplicating identifiers.
Worked example
- Input
- Cоnfig file: аdmin2, Noёl, café
- Scripts found
- Latin (Latn), Cyrillic (Cyrl)
- Characters flagged
- 7
| Index | Character | Code point | Script | Replacement | Rule |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | Latin | C | NFKC normalization |
| 1 | о | U+043E | Cyrillic | o | Confusable map |
| 7 | fi | U+FB01 | Latin | fi | NFKC normalization |
| 10 | : | U+FF1A | Common | : | Confusable map |
| 12 | а | U+0430 | Cyrillic | a | Confusable map |
| 17 | 2 | U+FF12 | Common | 2 | Confusable map |
| 22 | ё | U+0451 | Cyrillic | ë | Confusable map |
- Preserve Readable Unicode
- Config file: admin2, Noël, café
- Strict ASCII Fallback
- Config file: admin2, Noel, café
How it works
- Known look-alikes are replaced from the bundled map first. Characters outside the map are normalized with NFKC, which folds fullwidth forms, ligatures, and other compatibility characters into their standard equivalents.
- Preserve Readable Unicode keeps an accented letter where the map defines one, such as Cyrillic ё → ë. Strict ASCII Fallback uses the plain ASCII letter instead (ё → e).
- Letters that are neither in the map nor changed by NFKC are kept, so café keeps its accent in both modes. Normalized text is a comparison key, not a security verdict: store the original as well and review flagged characters.
Conversion is best-effort: mapped confusables and NFKC folding are deterministic, but some legitimate Unicode will not be flagged.
Your text
Paste or type — results update as you type (lightly debounced for long input).
30 characters scanned
7 suspicious
Strict ASCII Fallback
Original (suspicious characters marked)
Suspicious characters in the original view are underlined and labeled “susp.” in addition to highlight color.
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
Cleaned output
Character analysis
| Index (0-based) | Original | Replacement | Code point | Reason |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |