Detect homoglyphs in domains, usernames, and messages
A homoglyph is a character that looks like a different, more familiar one: Cyrillic а (U+0430) and Latin a (U+0061) are identical in most fonts but are different code points. The detector checks every character and lists each look-alike with its position, code point, Unicode script, and the character it imitates.
Worked example
- Input
- pаypаl.com / Αdmіn
- Scripts found
- Latin (Latn), Cyrillic (Cyrl), Greek (Grek)
- Characters flagged
- 4
| Index | Character | Code point | Script | Replacement | Rule |
|---|---|---|---|---|---|
| 1 | а | U+0430 | Cyrillic | a | Confusable map |
| 4 | а | U+0430 | Cyrillic | a | Confusable map |
| 13 | Α | U+0391 | Greek | A | Confusable map |
| 16 | і | U+0456 | Cyrillic | i | Confusable map |
- Preserve Readable Unicode
- paypal.com / Admin
How it works
- Every character is checked against a bundled map of common Cyrillic, Greek, Latin-variant, and fullwidth look-alikes. A match is flagged together with the ASCII letter, digit, or punctuation mark it imitates.
- Characters that are not in the map go through Unicode NFKC normalization. If NFKC changes them, as it does for fullwidth letters and ligatures such as fi, they are flagged as compatibility forms.
- Positions count Unicode code points from 0. A word that mixes Latin letters with Cyrillic or Greek ones is a strong warning sign, because legitimate words rarely switch scripts halfway through.
Conversion is best-effort: mapped confusables and NFKC folding are deterministic, but some legitimate Unicode will not be flagged.
Your text
Paste or type — results update as you type (lightly debounced for long input).
18 characters scanned
4 suspicious
Preserve Readable Unicode
Original (suspicious characters marked)
Suspicious characters in the original view are underlined and labeled “susp.” in addition to highlight color.
psuspicious character аypsuspicious character аl.com / suspicious character Αdmsuspicious character іn
Cleaned output
Character analysis
| Index (0-based) | Original | Replacement | Code point | Reason |
|---|---|---|---|---|
| 0 | p | p | U+0070 | Not flagged as a confusable or compatibility character. |
| 1 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | y | y | U+0079 | Not flagged as a confusable or compatibility character. |
| 3 | p | p | U+0070 | Not flagged as a confusable or compatibility character. |
| 4 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 5 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 6 | . | . | U+002E | Not flagged as a confusable or compatibility character. |
| 7 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 8 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 9 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 10 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 11 | / | / | U+002F | Not flagged as a confusable or compatibility character. |
| 12 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 13 | Α | A | U+0391 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 14 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 15 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 16 | і | i | U+0456 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 17 | n | n | U+006E | Not flagged as a confusable or compatibility character. |