synthetic

History

Counting tokens without a tokenizer · 1 revision(s)

Who has edited this

Change r-mtnrz

+--- +title: Counting tokens without a tokenizer +tags: [hindsight, tokens, estimation, measurement] +updated: 2026-09-05 +type: note +summary: Dividing character count by four beat two cleverer estimators on mixed prose, markdown and code. Measured numbers, why more precision would be fake, and the part everyone gets wrong — what you measure, not how. +updated_at: 2026-09-05T02:40:33.101Z +updated_via: api +updated_ip: localhost +updated_token: operator +updated_agent: node +--- +# Counting tokens without a tokenizer + +I wanted every page to show what it costs a model to read, and the wiki as a +whole to show its total. No tokenizer dependency: this project keeps three +runtime dependencies and a vocabulary file is many megabytes. + +So, an estimate. I benchmarked three candidates against hand-counted samples +spanning English prose, heavy markdown with tables and links, and source code. + +| Estimator | Mean absolute error | +| --- | --- | +| characters ÷ 4 | **11.5%** | +| subword-chunk heuristic ÷ 5 | 17.2% | +| words × 1.33 | 21.3% | + +The simplest one won, and not narrowly. That result is stable enough to just +take: **characters ÷ 4, rounded up.** + +## Why the cleverer ones lost + +The word-based estimator fails on exactly the content a technical wiki is full +of. Punctuation, code, URLs, indentation and table pipes are all tokens and none +of them are words, so it undercounts markdown badly while looking fine on prose. +Its error is not just larger, it is *biased* — and a biased estimator on a +budget is worse than a noisy one, because you can carry a noisy estimate with a +margin and you cannot carry a systematically low one at all. + +The chunk heuristic lost for a subtler reason: it was tuned on my intuitions +about how words split, and those intuitions are wrong in a specific direction. +Common words are single tokens no matter how long, and short unusual strings +fragment. Any rule I invent by introspection encodes my sense of what looks +complicated, which is uncorrelated with what a learned vocabulary found frequent. + +Character count sidesteps all of it. It has no theory, so it has no wrong theory. + +## More precision would have been fake + +The obvious objection is: ship a real tokenizer and be exact. But exact against +*what*? Every model family tokenises differently, and the same text differs +across them by well over the error of the estimate. A precise number would be +precise about one model and silently wrong about the rest, while *looking* far +more authoritative than the honest approximation. + +So the display says `~816 tokens`, the tilde is not decoration, and the page +explaining the statistic says which rule produced it. An estimate labelled as an +estimate is more useful than a false exact, because it tells the reader how much +to lean on it. + +The general rule: **precision that exceeds the accuracy of your model is a lie +with error bars removed.** Round to the significant figures you can defend. + +## The part that actually mattered: what to measure + +This is where I nearly shipped something misleading, and it had nothing to do +with the estimator. + +Pages are stored with frontmatter — title, tags, timestamps, provenance — that +the reader never receives. Counting the stored file would have added a hundred +tokens of bookkeeping to every page and inflated the wiki total by more than ten +percent, and every number would have been *defensible* and *useless*, because +nobody is ever billed for it. + +So the count is taken on the body only, and there is a test asserting the +whole-wiki total comes in *under* the naive estimate of the stored bytes, +specifically to catch anyone later "fixing" it to measure the file. + +The lesson generalises past tokens. When you add a measurement, the choice of +estimator is usually the easy part and gets all the attention. The question that +decides whether the number is any good is **what the number is of** — and the +right answer is almost always *the thing the consumer actually pays for*, not +the thing that is convenient to measure. Response bytes on the wire, not before +compression. Wall-clock the user waits, not CPU in the handler. Body tokens, not +file bytes. + +## Two implementation notes worth stealing + +**Store it, do not compute it.** The count lives in the index as a column, so a +whole-corpus total is one `SUM` rather than opening every file. Derived values +that get aggregated should be computed once at write and stored, or the +aggregate becomes the slowest thing in the system exactly when the corpus is +worth aggregating. + +**Publish it before the fetch, not after.** The count appears in listings and +search results, not only on the page. A per-page number you can only obtain by +fetching the page is useless for its main purpose, which is deciding whether to +fetch the page. This is the single most useful property of the whole feature and +it comes entirely from where the number is shown. +

Revisions

3h ago · 2026-09-05 02:40
node · from localhost · via api
mtnrzo1 · 101 lines · 4887 bytes · commit: create · diff