Skip to content

Use this result well

Inputs that matter
Text typed, pasted or imported from a plain text file; up to 500,000 UTF-16 code units.
Output to expect
Named Unicode counts, word frequencies and timing estimates, with English readability and a publishing-field comparison where shown.
How it works
Analysis updates as you edit; larger drafts run in a cancellable background task. Reading uses 238 words per minute and speaking uses 150.
  • Character counts distinguish UTF-16 units, Unicode code points and visible character groups. Use the measure required by your destination.
  • Word and sentence boundaries can vary by language and browser. English syllable and readability estimates do not judge accuracy or suitability for your audience.
  • Select a publishing field for its named character or byte comparison. An SEO writing budget is your own editing target.
Was this tool helpful?

Reference & details

How it works

Words and boundaries

Native Intl.Segmenter identifies word-like units. Hyphenated expressions can split into several tokens. If native segmentation is unavailable, the tool discloses its CJK-character and whitespace approximation. Blank lines separate paragraph blocks.

Fixed timing estimates

Reading uses 238 words per minute, an adult English non-fiction reference from Brysbaert (2019). Speaking uses a 150-word-per-minute planning assumption. These rates do not measure your own pace or include pauses.

Reading minutes = counted words / 238; speaking minutes = counted words / 150

Source: Brysbaert (2019): reading-rate review

English readability and publishing fields

The English prose control enables inspectable formula estimates and word pronunciation corrections. Publishing comparisons use the selected field and named Unicode or byte unit. An SEO budget is an editing target you enter yourself.

Updated: September 2026

Example Scenarios

Paste The cat sat on the mat. to inspect six word-like tokens and one sentence candidate.

Text: The cat sat on the mat.

6 words; 1 sentence candidate; 17 letters and numbers for ARI.

Keep the blank line between the two sentences.

Text: One two three. Four five six.

6 word-like tokens, 2 sentence candidates and 2 paragraph blocks.

Find repeated content words without changing the draft.

Text: constructor constructor prototype.

constructor occurs 2 times; prototype occurs 1 time.

FAQ

The report distinguishes UTF-16 code units, Unicode code points and grapheme clusters. They can differ for emoji and combining marks. A selected publishing comparison can also count UTF-8 bytes or X weighted characters.

Word-like tokens are normalized to NFC, lowercased and stripped of non-letter, non-number and non-mark characters for frequency keys. The panel shows the ten most frequent tokens, with an optional view excluding its listed common English words. All counted tokens contribute to the denominator.

No. This workspace reports on the supplied text. It does not transform CSV columns, repair grammar or format source code. Copy Stats copies a statistics report rather than a rewritten draft.

No. Each comparison names the field, counting unit and source. The YouTube Data API description uses bytes; X uses weighted characters and recognized URL spans. For sources that do not specify a Unicode unit, compare the displayed range and verify the actual composer.

Paste extracted text or choose a plain text file. Files may use UTF-8, or UTF-16 with an encoding marker. The limit is 2,000,000 file bytes and 500,000 UTF-16 code units after decoding. PDF and Word document parsing is not provided.

Larger inputs run in a browser worker. You can cancel or retry the analysis, or edit the source to replace it. Copy, Print and Project saves wait for a result that matches the current input.

The report names each counting unit and the reading and speaking rates. Where shown, it also includes readability formulas, their counted inputs, word corrections, sources and the selected publishing-length comparison.

About Text Statistics

Inspect the same draft across named character units, word frequencies, sentence candidates, timing estimates and English readability formulas. Choose a destination field for a separate length comparison, then copy or print a report when the current analysis is ready.