ToolForge
Advertisement

Word Frequency Counter

Count words, phrases or letters, with context and comparison

Written by toolforge.websiteLast reviewed How we build and check these tools

Word Frequency Counter tool

Waiting for text.

Your text
Waiting for text

Ready-made settings

Count

One row per word, however you define a word below.

What counts as a word

Off, Apple and apple count together and the row shows the spelling you used most.

don't stays one word. Curly apostrophes count as the same character.

state-of-the-art as a single term rather than four.

Off, 2024 and 10 are not counted as words.

Tags, code fences and link targets stop counting as words.

One long URL otherwise contributes a dozen nonsense terms.

Suffix rules only: runs and running join run; ran does not. Plurals, possessives, -ed and -ing.

Words to ignore
Showing the results

Results

Paste some text above and the table appears as you type.

Every count happens in this browser. Nothing you paste is uploaded, and the text is only stored on this device while “Keep my text in this browser” is ticked.

Word Frequency Counter: key facts

What it does
Count words, phrases or letters, with context and comparison
Category
Text Tools
Cost
Free, with no account, sign-up, or install.
Your data
Runs entirely in your browser — the files and text you enter are never uploaded to a server.
Last reviewed
. Report an incorrect result.
Advertisement

What the Word Frequency Counter does

A frequency table is the fastest way to find out what a piece of writing is actually about, and the fastest way to catch the word you have leaned on eleven times without noticing. This counter reads whatever you paste and ranks every term in it, with the count, the share of the text it occupies, and how far in it first shows up.

The thing that makes one frequency count differ from another is the definition of a word, so that definition is yours here rather than buried in a regular expression. Capitals can be folded together or kept apart. Hyphenated compounds can be one term or four. Apostrophes can stay inside contractions or be treated as breaks. Numbers can join the tally or sit it out. Markup and links can be stripped before anything is counted, and English plurals, possessives and -ed and -ing endings can optionally be folded onto a common stem so that "design", "designs" and "designed" stop competing with each other for the top of the list.

Words are not the only unit. Switch to phrases and it counts runs of two to five adjacent words — the pairs and triples that reveal a stock sentence or a habit of phrasing — and it refuses to build a phrase across a full stop, a line break or a bracket, so nothing in the table is an accident of two neighbouring sentences. Switch to letters and it counts single characters, which is what a cryptogram, a word game or a keyboard-layout argument needs.

The list of words to leave out is visible and editable in English, Spanish, French and German, because a filter that silently decides most of the table should not be a mystery. Click any row and every occurrence appears with the words on either side, so a count turns into an edit you can actually make. And a second box lets you count another text with identical settings and see, term by term, which words the two use differently.

Phrases, letters, and reading a term in context

Phrase counting is where a repetition problem usually becomes visible, because writers repeat constructions more often than they repeat single words. Phrases are assembled from the untouched sequence of words and only then judged against the ignore list, which is deliberate: dropping filler first would splice non-adjacent words together and report pairs that never appeared on the page. You choose whether phrases containing an ignored word are kept, dropped when one sits at either end, or dropped outright.

Letter mode counts single characters, either letters alone or every character that is not a space. It is the mode for solving a substitution cipher, checking the letter distribution of a puzzle, or settling an argument about which keys a language actually uses.

Clicking a row opens its concordance — up to two hundred occurrences, each with about fifty characters either side and the term itself marked, prefixed by how far through the text it sits. Seeing the same word in twelve contexts at once is what tells you whether it is a tic worth cutting or a technical term doing its job. The word cloud is the same data drawn loosely: the top sixty terms sized by the square root of their counts, so the most common one does not swallow the panel.

Using the Word Frequency Counter, step by step

  1. Paste your text into the box, or drop a .txt, .md, .html or .csv file onto it.
  2. Choose what to count: single words, phrases of two to five words, or letters.
  3. Under "What counts as a word", set case, hyphens, apostrophes, numbers, minimum length, and whether markup and links come off first.
  4. Pick a list of words to ignore and edit it, or type your own into "Also ignore these".
  5. Sort, search, or raise the minimum count, then click any row to read that term in context.
  6. Copy or download the table as CSV, TSV, Markdown or JSON, or save the plain-text report.

What the percentage is a percentage of

Every share in the table is measured against the terms that were actually counted, and the figure it is measured against is printed above the table: so many found, so many counted, so many set aside. That distinction matters the moment you switch an ignore list on. If a hundred-word passage contains forty function words and you ask for them to be left out, the remaining sixty are the population, and the column adds up to a hundred percent of that sixty rather than to sixty percent of something you can no longer see.

Where the filters have removed anything, a second column appears giving each term's share of everything found, filters included. The two readings answer different questions — one is "how much of the writing that carries meaning is this word", the other is "how much of the document is this word" — and neither is more correct, so both are on offer rather than one being chosen for you.

Shares are printed with enough precision to stay useful in a long document. In a hundred thousand words, a term used once and a term used four times would both round away to nothing at one decimal place, so small values are given three.

  • 1,204 counted · 809 set aside — the second number is what the ignore list, the number rule and the minimum length removed.
  • A term used four times in 100,000 words reads 0.004%, not 0.0%.
  • Turning the ignore list off changes every percentage, because it changes the population.

What makes this one worth using

  • Words, two- to five-word phrases and single letters, all from the same settings, so you can move between them without re-entering anything.
  • Text is read as Unicode, so accented words, Cyrillic, Greek and other scripts are counted properly rather than silently mangled or dropped.
  • Percentages are shares of what was actually counted, and the number they are shares of is printed above the table.
  • The list of words to leave out is visible, editable and available in four languages, instead of being a hidden decision about your results.
  • Clicking a term shows every occurrence with the words around it, which is what makes the table something you can act on.
  • A second box counts another text with identical settings and reports the difference in each term's share, in percentage points.
  • Everything you see can leave: copy as CSV, TSV, Markdown or JSON, or download the table or a plain-text report.
  • All of it is computed in your browser — no upload, and the text is only kept on your own device if you tick the box.

Counting spellings is not counting meanings

A tally works on the characters in front of it. Unless you ask otherwise, "walk", "walks", "walked" and "walking" are four entries describing one idea, and a word that carries two meanings — a company and a fruit sharing the spelling "apple" — is one entry describing two. This is a limitation of counting rather than a fault in any particular counter, and it is worth holding in mind before treating a ranking as a summary.

Two controls narrow the gap. Grouping word forms applies English suffix rules, folding plurals, possessives and -ed and -ing endings onto a stem and undoing a doubled consonant so that "running" reaches "run". It is rules, not a dictionary, which means it will never turn "ran" into "run" or "better" into "good"; the control says so, it is off unless you switch it on, and every spelling that was merged is listed on the row so you can see what happened. Turning capitals back on goes the other way, separating a proper noun from the ordinary word it shares its letters with.

There is one more habit worth resisting: reading frequency as importance. A term at the centre of a document may appear rarely because the writer worked to avoid repeating it, while a term that recurs constantly may be scaffolding. The standard fix weighs a word by how unusual it is across a whole collection of documents, which needs a collection. The comparison box is the version of that idea you can run without one — count a second text and the change column shows which words this text leans on that the other does not.

Frequently Asked Questions

What exactly counts as one word?

A run of letters, digits and accent marks, optionally joined on the inside by an apostrophe or a hyphen. So "don't" is one word by default, quotation marks around 'hello' are not part of it, and an underscore separates rather than joins. Whether hyphenated compounds count as one term or several is a setting, as is whether pure numbers like 2024 are counted at all.

Why do the percentages change when I turn the ignore list on?

Because the population changes. Shares are measured against the terms actually counted, so removing filler words from the tally makes the remaining words a larger part of a smaller whole. The counted total and the number set aside are both shown above the table, and where anything has been filtered a second column gives each term's share of everything found.

Can it treat "run", "runs" and "running" as the same word?

Yes — switch on "Group English word forms". It folds plurals, possessives and -ed and -ing endings onto a common stem and undoes a doubled consonant, so "running" joins "run". It is a set of suffix rules rather than a dictionary, so it will not connect irregular forms like "ran" or "went", and each row lists the spellings that were merged into it.

How do I find the phrases I repeat, not just the words?

Choose Phrases and set the length to two or three. Phrases are built from adjacent words and never span a full stop, a line break or a bracket, so what you see genuinely appeared that way. Setting the minimum count to 2 hides everything that occurred only once and leaves the repeated constructions.

How is this different from a keyword density checker?

A density checker judges a page against a target keyword and a density range for search engines. This one has no target and no ideal range: it describes what the text contains, in words, phrases or letters, with context lines, word-form grouping and a two-text comparison. Use the density checker when you are optimising a page, and this when you are reading or editing one.

Is anything sent anywhere?

No. The tokenizing, counting, comparison and every export happen in your browser, and no part of your text is transmitted. If you leave "Keep my text in this browser" ticked, the text and your settings are saved in this browser's local storage on this device only; untick it and they are removed. Very long texts are never written to storage at all.

Related Tools

Advertisement
Buy Me a Coffee