A word frequency counter answers which tokens appear, and how often. It is not a total. The Word Counter reports words, characters, sentences, and paragraphs for the whole paste. It does not rank the pieces. This page splits the paste, counts each remaining token, and lists them by count.
A token is a piece separated by Unicode whitespace. On each piece, leading and trailing characters in the Unicode punctuation class are removed. Characters inside the piece stay, so don’t and well-known remain one token each when the apostrophe or hyphen sits in the middle. A piece that is only punctuation is discarded. Empty pieces are discarded.
Match case is off unless you turn it on. Off means the token is folded with JavaScript’s ordinary toLowerCase before it is counted. That conversion is not a Turkish locale fold. On means A and a are two tokens. The page does not stem words, detect a language, drop a stop-word list, or group synonyms. New York is two tokens because of the space.
The table shows at most 500 rows. Those rows are the first 500 after sorting. The token total and the distinct total always include every token, including ones left off the table. The status says how many distinct tokens were omitted when the cap is hit. Do not treat the visible rows as the whole set when that note is present.
The example text is “One, one; two.” With Match case off, the comma, semicolon, and period are stripped, then case is folded. The table has one with count 2 and percent 66.7, then two with count 1 and percent 33.3. With Match case on, One and one would be separate rows.
Three different tokens that each appear once are each 33.3 percent. Those rounded figures add to 99.9, not 100.0. The page does not adjust the last row to force a total of 100. One token out of eighty is 1.25 percent before rounding and displays as 1.3, because halves round away from zero.
One, one; two.
one 2 66.7
two 1 33.3
Sort order is count descending. When two tokens share a count, they are ordered by UTF-16 code units ascending. The page does not use a locale sort, so the order stays the same on every browser. A token that begins with an earlier code unit comes first. That can differ from a dictionary order that groups accented letters with the unaccented letter.
The count is a mechanical split. It is not a publisher’s house style, not a search-engine score, and not a reading of what the text means. If you only need the length of the paste, the word counter is the shorter page. If you need to delete invisible marks inside a word, that is a different tool and a different character set.
Counting runs in the browser. The paste is not uploaded. Clearing the form removes it from the field. Do not paste passwords, customer lists, or unpublished material you would not want left on the screen.
| Token | Count | Percent |
|---|
Tokens are whitespace-separated pieces. Leading and trailing punctuation is removed. Match case is off by default and folds with the ordinary JavaScript lower-case conversion, which is not a Turkish locale fold. The table shows at most 500 rows. Totals include every token.
Each percent is rounded to one decimal place on its own, half away from zero. The page does not change the last row to force a sum of 100.0.
No. Folding uses JavaScript toLowerCase with no locale argument. A dotted capital I is not given a special Turkish mapping.
The table shows the first 500 rows in the documented sort. The token total and the distinct total still include the rows that are not drawn.
Yes, when the apostrophe is internal. Only leading and trailing punctuation is stripped. A space still splits tokens.