Skip to content

Character Frequency Analyzer

Analyze Unicode scalar or grapheme frequencies with explicit text policies, correct percentages, invisible-character review and complete JSON reports.

Choose your path

Built around the job you need to finish

Count Unicode scalars or grapheme clusters with explicit normalization, case and whitespace policies; inspect every frequency and export the full distribution.

Localization engineer

Distinguish emoji and combining sequences from UTF-16 units.

Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.

Counts and percentages share the selected symbol denominator.

Text-data reviewer

Know whether normalization and whitespace change a comparison.

Compare NFC, lowercase and whitespace settings while retaining original text.

Every transformation is explicit and no invisible format character silently disappears.

Mobile corpus inspector

Review many distinct symbols without losing the complete distribution.

Import UTF-8 with BOM/CRLF, page 220 unique symbols, copy/export and undo replacement.

Preview limits do not truncate the report or destroy the source.

Authoritative checks for this tool

Outputs and checklists are planning aids. Review the linked current authorities and the records, terms, instructions, and requirements that apply to your exact situation before a consequential decision.

Was this tool helpful?

Reference & details

How it works

Counting basis

Choose Unicode scalars or runtime grapheme clusters. UTF-16 source length is separately labeled; frequency percentages divide by the selected symbol total.

Explicit text policies

Optional NFC and default lowercase precede whitespace filtering. These are analysis choices, not edits to the source or full linguistic case folding.

Complete review

Review up to 200 rows, 20 per page. The JSON report includes every result row, settings and methods. Code points identify hidden characters.

Updated: August 2026

Example Scenarios

Distinguish emoji and combining sequences from UTF-16 units. Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.

Know whether normalization and whitespace change a comparison. Compare NFC, lowercase and whitespace settings while retaining original text.

Review many distinct symbols without losing the complete distribution. Import UTF-8 with BOM/CRLF, page 220 unique symbols, copy/export and undo replacement.

Common Mistakes to Avoid

Using UTF-16 length as the frequency denominator

Use the selected symbol total and compare separately labeled source units.

Treating normalization as harmless for every task

Keep exact source data and choose normalization only when the comparison calls for it.

FAQ

JavaScript UTF-16 length and Unicode scalar count differ. This workspace labels them separately and uses one consistent denominator for frequencies.

They are runtime segmentation clusters, not a universal promise about rendering or language. Inspect the actual text in its target environment.

No. NFC applies only to the comparison stream. The editor and original file remain your source of record.

Unicode White_Space scalars are removed before segmentation. BOM and zero-width format characters are not included in that property.

Yes. Copy or download the complete JSON report; the 200-row display limit does not cap report rows. Check that the browser actually saved the file.

About Character Frequency Analyzer

Count Unicode scalars or grapheme clusters with explicit normalization, case and whitespace policies; inspect every frequency and export the full distribution.