Localization engineer
Distinguish emoji and combining sequences from UTF-16 units.
Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.
Counts and percentages share the selected symbol denominator.
Analyze Unicode scalar or grapheme frequencies with explicit text policies, correct percentages, invisible-character review and complete JSON reports.
Choose your path
Count Unicode scalars or grapheme clusters with explicit normalization, case and whitespace policies; inspect every frequency and export the full distribution.
Distinguish emoji and combining sequences from UTF-16 units.
Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.
Counts and percentages share the selected symbol denominator.
Know whether normalization and whitespace change a comparison.
Compare NFC, lowercase and whitespace settings while retaining original text.
Every transformation is explicit and no invisible format character silently disappears.
Review many distinct symbols without losing the complete distribution.
Import UTF-8 with BOM/CRLF, page 220 unique symbols, copy/export and undo replacement.
Preview limits do not truncate the report or destroy the source.
Outputs and checklists are planning aids. Review the linked current authorities and the records, terms, instructions, and requirements that apply to your exact situation before a consequential decision.
Tools you might need next
Compare UTF-16 units, Unicode code points and grapheme clusters. Choose a publishing field to inspect its weighted-character, Unicode or UTF-8 byte rule.
Browse every word and 2- or 3-word phrase frequency, search and filter rows, inspect source matches, and export complete Unicode-aware analysis.
Find duplicate line groups with exact source positions and original values. Choose line, case and whitespace policies and export a complete review report.
Choose Unicode scalars or runtime grapheme clusters. UTF-16 source length is separately labeled; frequency percentages divide by the selected symbol total.
Optional NFC and default lowercase precede whitespace filtering. These are analysis choices, not edits to the source or full linguistic case folding.
Review up to 200 rows, 20 per page. The JSON report includes every result row, settings and methods. Code points identify hidden characters.
Updated: August 2026
Distinguish emoji and combining sequences from UTF-16 units. Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.
Know whether normalization and whitespace change a comparison. Compare NFC, lowercase and whitespace settings while retaining original text.
Review many distinct symbols without losing the complete distribution. Import UTF-8 with BOM/CRLF, page 220 unique symbols, copy/export and undo replacement.
Use the selected symbol total and compare separately labeled source units.
Keep exact source data and choose normalization only when the comparison calls for it.
Count Unicode scalars or grapheme clusters with explicit normalization, case and whitespace policies; inspect every frequency and export the full distribution.