Verwarrende tekens normaliseren naar veilige tekst
Bij het normaliseren van verwarrende tekens wordt elke dubbelganger vervangen door het gewone teken dat hij nabootst, zodat twee tekenreeksen die er hetzelfde uitzien ook bij vergelijking gelijk zijn. Doe dit voordat je gebruikersnamen vergelijkt, blokkeerlijsten toepast of dubbele ID's verwijdert.
Uitgewerkt voorbeeld
- Invoer
- Cоnfig file: аdmin2, Noёl, café
- Gevonden schriftsystemen
- Latijns (Latn), Cyrillisch (Cyrl)
- Gemarkeerde tekens
- 7
| Positie | Teken | Codepunt | Schrift | Vervanging | Regel |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | Latijns | C | NFKC-normalisatie |
| 1 | о | U+043E | Cyrillisch | o | Verwarringstabel |
| 7 | fi | U+FB01 | Latijns | fi | NFKC-normalisatie |
| 10 | : | U+FF1A | algemeen | : | Verwarringstabel |
| 12 | а | U+0430 | Cyrillisch | a | Verwarringstabel |
| 17 | 2 | U+FF12 | algemeen | 2 | Verwarringstabel |
| 22 | ё | U+0451 | Cyrillisch | ë | Verwarringstabel |
- Behoud leesbare Unicode
- Config file: admin2, Noël, café
- Strikte ASCII-terugval
- Config file: admin2, Noel, café
Zo werkt het
- Bekende dubbelgangers worden eerst vervangen volgens de ingebouwde tabel. Tekens buiten de tabel worden genormaliseerd met NFKC, dat volbreedtevormen, ligaturen en andere compatibiliteitstekens omzet naar hun standaardequivalenten.
- Behoud leesbare Unicode houdt een letter met accent aan als de tabel er een definieert, zoals Cyrillische ё → ë. Strikte ASCII-terugval gebruikt in plaats daarvan de gewone ASCII-letter (ё → e).
- Letters die niet in de tabel staan en niet door NFKC worden veranderd, blijven behouden, dus café houdt in beide modi zijn accent. Genormaliseerde tekst is een vergelijkingssleutel, geen beveiligingsoordeel: bewaar ook het origineel en controleer gemarkeerde tekens.
Conversie is de beste inspanning: in kaart gebrachte confusables en NFKC-vouwing zijn deterministisch, maar sommige legitieme Unicode wordt niet gemarkeerd.
Jouw tekst
Plakken of typen: de resultaten worden bijgewerkt terwijl u typt (licht gedebounced voor lange invoer).
30 tekens gescand
7 verdacht
Strikte ASCII-terugval
Origineel (verdachte tekens gemarkeerd)
Verdachte tekens in de oorspronkelijke weergave zijn onderstreept en voorzien van het label 'verdacht'. naast de accentkleur.
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
Opgeschoonde uitvoer
Tekenanalyse
| Index (0-gebaseerd) | Origineel | Vervanging | Codepunt | Reden |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |