Skip to content

Use this result well

Inputs that matter
Source text up to 500,000 UTF-16 units and a separate optional keyword or phrase up to 2,000 units, with match type, case, overlap and word-language controls.
Output to expect
Occurrences and their denominator, selectable source locations, searchable 1-, 2- and 3-word frequency tables, filtered CSV and a complete frequency-analysis JSON report.
How it works
Whole-word matching follows browser word-like boundaries. NFC and Unicode 17 default full case folding preserve accents and internal punctuation. A phrase cannot cross punctuation between words; text fragments must span complete source graphemes.
  • Whole-word rate = occurrences / all source words × 100. Text-fragment rate uses source grapheme clusters; distinct word coverage is reported separately.
  • Frequency groups always count overlapping windows. Hiding common English words filters rows without changing the source denominator or creating new phrase neighbors.
  • Match locations preview the first 50. Browse every frequency row in pages of 50; JSON retains every row and the exact source. Project summaries retain the first 10 single-word rows.
  • Counts help inspect repetition in context. They do not give an optimal SEO keyword density, semantic importance or ranking prediction.

Choose your path

Built around the job you need to finish

Count and rank the top displayed NFC/case-normalized word-like tokens in the exact entered text with explicit segmentation, tie, truncation and semantic limits.

Content editor checking repetition

See repeated surface tokens without a hidden stop-word or importance claim.

Enter the exact draft, inspect total/unique tokens and top rows, then review repetitions in context.

Can revise or retain wording without interpreting frequency as quality or engagement.

Multilingual or corpus reviewer

Preserve accented and non-Latin letters and know how word-like boundaries were derived.

Test case, accents, combining marks, CJK and mixed scripts; inspect NFC normalization and runtime segmentation disclosure.

Can reproduce displayed rows and identify linguistic limitations.

Claims, brand or accessibility reviewer

Prevent frequency output from becoming a sentiment, originality, keyword, bias or clarity verdict.

Compare the rows with the claim/source matrix and affected-person review rather than optimizing automatically.

Can reject unsupported interpretation while retaining the descriptive record.

Was this tool helpful?

Reference & details

How it works

Word boundaries

Choose the word language used by browser Intl.Segmenter. Browser dictionary versions can change boundaries. When word segmentation is unavailable, the disclosed Unicode-run fallback preserves letters, marks, numbers and internal apostrophes but does not provide dictionary segmentation.

Source: ICU Boundary Analysis

Normalization and case

Case-insensitive counts use NFC, Unicode 17.0.0 default full case folding, then NFC again. Accents and internal punctuation remain. Case-sensitive mode still applies NFC; word language does not select locale-tailored casing. There is no stemming or synonym grouping.

Source: Unicode 17.0.0 CaseFolding.txt

Adjacent phrase groups

Tables count overlapping windows of 1, 2 or 3 words. Words can be separated by whitespace or touch directly; punctuation between tokens breaks a phrase. Every rate uses all source word-like tokens as its denominator. Hiding common English words filters rows after counting.

Complete rows and inspectable summaries

Rows sort by descending count, then normalized UTF-16 text order for ties. Search, sorting and pages of 50 reach every row. Filtered CSV includes all matching rows; JSON includes every group and the exact source. Copy summarizes the first 20 rows per group; Project saves the first 10 unfiltered single-word rows.

Interpret repetition in context

These rates describe the entered text and do not give an optimal keyword density or predict rankings. Google Search guidance describes unnatural keyword repetition as keyword stuffing. Review whether each occurrence helps the reader.

Source: Google Search spam policies: keyword stuffing

Updated: September 2026

Example Scenarios

Select 2-word phrases and inspect a frequent row in the source. Match that row to examine more locations before editing the passage.

Café and CAFÉ share an NFC form. Straße and STRASSE group with case sensitivity off. Review word-language settings and the source when interpreting mixed-script copy.

Search or hide common words to focus a CSV export, then download analysis JSON to keep all counts, options and the exact original text. A Project summary retains the first 10 single-word rows.

FAQ

Yes. Browse all rows in pages of 50, search for a row, or download every matching row as CSV. JSON retains all 1-, 2- and 3-word groups. The first-10 limit applies only to Project summaries.

No stemming, lemmatization or synonym grouping is performed. Case-insensitive mode uses Unicode default full case folding without locale-tailored casing; accents and internal punctuation remain. Word language changes segmentation, not the case-fold rules.

It hides any frequency row containing a word from the disclosed list, including phrase rows. It does not remove words from the source, join new neighbors, change the rate denominator or alter the complete JSON report.

Source selects the first location of a row. Match uses that word or phrase as a query, then the matching-locations section lets you select up to 50 locations in the original editor. Counts cover the whole source.

Frequency alone does not establish importance, originality, truth, clarity or search performance. Repeated names, function words and required disclosures can be appropriate. Review the term in its context.

About Unicode Word Frequency Counter

Find repeated words and phrases in your exact source text. Browse all frequency rows, filter common English words, or use a row as a keyword to inspect its matching locations. Counts describe surface forms, so context still determines what to keep or change.