Roadmap and API boundaries¶
The order below keeps performance work evidence-driven and prevents provider catalog concerns from leaking into the exact tokenizer kernel.
See the provider capability matrix for the current per-provider status of exact counting, estimated counting, and model-based cost estimation.
Exact-counting performance¶
- Implemented tokenizer definitions:
o200k_base,cl100k_base,p50k_base, andr50k_base. - Implemented strict UTF-8 Binary input beside String, Categorical, and Enum. Only non-null binary values are validated; see Binary input.
- Establish controlled-host throughput, scaling, CPU, and peak-memory baselines with the benchmark matrix.
- Profile pre-tokenization, BPE merging, output construction, and Polars integration independently.
- The Criterion harness now pairs direct count-only and token-ID encoding on identical strings; this measures their public-operation difference, not a clean separation of allocation and BPE work.
- A native component benchmark now separates borrowed-string counting,
StringChunkedtraversal, Arrow output construction, and output-only work across 24-, 128-, and 2,048-byte rows. It is a coarse comparison; pre-tokenization versus BPE still needs a symbol-resolving profiler or explicit tokenizer instrumentation. - Implemented: categorical/enum inputs count each used dictionary value once with dense, sparse, and parallel physical-ID paths.
- Opt-in bounded whole-value caching is available for repeated String or Binary text. Experimental benchmarks cover 0.01%, 0.1%, 1%, 10%, 50%, and 100% cardinality. Repeat them on a controlled host with end-to-end memory measurements before choosing an automatic policy.
- The benchmark matrix now supports deterministic exceptionally large rows. A dominant-row guard skips parallel dispatch when less than 512 KiB of independent text remains; continue profiling huge rows without nested thread pools or changing exact results.
Token estimation¶
- Freeze a multilingual, multi-format training and held-out corpus.
- Implemented a benchmark-only error summarizer with zero-token handling and language/content/length/count/ASCII strata. It does not supply the corpus or a fitted estimator.
- Implemented a candidate-corpus validator and order-independent SHA-256 manifest. It rejects duplicate IDs and related or identical text crossing train/held-out splits.
- Added a reproducible, checksum-pinned Common Voice CC0 sentence candidate: 3,200 prose records in 10 languages across train/held-out splits. Multi-format coverage, translation-family auditing, and representativeness remain open.
- Added versioned
tiktokenexact-count labels for both supported encodings on that candidate, with zero local-kernel parity mismatches. No estimator is fitted or accuracy claim made. - Added a pinned MIT-licensed JSON Schema fixture candidate: 46 intact files, including 35 long records, that can be combined with the prose candidate. These conformance fixtures are not representative application JSON; code, broader structured text, and coverage balancing remain open.
- Added 49 pinned Python source files, with Flask in train and Black held out to avoid same-project leakage. The three-format candidate is still highly prose-weighted and not representative of other programming languages.
- Added a split-by-format coverage audit tied to corpus and oracle hashes. Sparse long-form and non-prose strata are explicit; sampling balance and external validity remain gates before estimator accuracy publication.
- Added a benchmark-only, train-fitted nonnegative ASCII/non-ASCII byte model compared against bytes/4 on held-out rows. It improves aggregate candidate MAE but has large JSON and code errors; production estimator work still requires CPU-cost measurements and broader external validation.
- Measure one-pass UTF-8 features by accuracy contribution per CPU cost.
- Fit deterministic, dependency-free estimators independently for each tokenizer family.
- Publish error distributions by language, content, length, token count, and ASCII/non-ASCII class.
Model aliases¶
A versioned model registry may map a provider-facing model identifier to a tokenizer definition. The low-level tokenizer continues to accept tokenizer identifiers only. Model aliases live in the Python-facing or a separate higher-level Rust layer.
Conceptually:
Implemented: the Python-facing API has a dated registry of exact OpenAI model
aliases for the locally supported tokenizers. count(model="gpt-5") resolves
to o200k_base before crossing the plugin boundary. The registry intentionally
does not use prefix matching: a newly released model must be reviewed and added
to a new registry version rather than silently inheriting an older mapping.
The initial mappings reproduce the supported subset of the pinned tiktoken
0.12.0 registry without importing tiktoken at runtime.
Gemini local tokenizers (investigation TODO)¶
Google's experimental Python LocalTokenizer currently gives us a promising
reference path, but not yet a blanket exactness guarantee for every Gemini
model. Its Gemma 3 path loads a hash-pinned SentencePiece model. Its
count_tokens() implementation then calls encode() and sums the lengths of
the returned token-ID lists, so a Polars-native count-only kernel may avoid
substantial allocation and Python/API overhead on large columns. Newer mapped
models use a separate Gemma 4 Hugging Face tokenizer path and must be evaluated
independently.
The pinned Gemma 3 artifact has now been inspected and is a SentencePiece
BPE model, not a Unigram model. See
the prototype notes for its validated configuration and
the reproducible oracle workflow.
Before adding public Gemini support:
- [ ] Benchmark Google's local tokenizer as shipped, its underlying
SentencePiece batch operation, and remote
countTokensseparately. Record initialization/download time, cold and warm throughput, peak RSS, and allocations for scalar and batch workloads. - [ ] Treat
gemma3/gemma4as versioned tokenizer definitions and Gemini names as model aliases. Do not put Gemini-specific branching in the tokenizer kernel. - [x] Prototype a pure-Rust, per-value, count-only SentencePiece BPE kernel
for the pinned Gemma 3 model without materializing token IDs. The prototype
is isolated behind the
gemma3-prototypeCargo feature and is not yet a public Polars expression. - [x] Preserve the pinned model's identity normalization, user-defined-symbol matching, ASCII-space escaping, and byte-fallback behavior exactly; do not apply an independent Unicode normalization pass. A 20,013-string initial oracle (3,321,397 bytes, 1,663,237 tokens) matched Google's local implementation.
- [ ] Compare against the official local implementation on the complete
multilingual/fuzz corpus and against remote
countTokensfor stable model IDs. Publish any raw-text versus request-accounting differences. - [ ] Reuse the existing Arrow, categorical/enum, byte-balanced parallel, and bounded-cache paths only after single-value parity is proven.
- [ ] Audit tokenizer artifact licensing and distribution. Pin the artifact URL, SHA-256, algorithm/configuration, and model-alias registry version; use an explicit verified download/cache flow if bundling is not permitted.
- [ ] Evaluate Gemma 4 separately, including its tokenizer artifact, processor behavior, dependency footprint, and whether a pure-Rust compatible path is possible.
- [ ] Add
gemma3orgemma4tocount()only after exactness is demonstrated for a pinned definition. If remote behavior cannot be reproduced locally, keep that mapping experimental or offer estimation rather than label it exact.
The initial scope remains raw text. Multimodal inputs, roles, tools, response schemas, and provider request serialization belong to future request-level accounting even where Google's local helper accepts some of those structures.
Reference implementations and definitions:
Model-based cost estimation¶
Cost estimation is a higher-level analytics feature and will accept a model, not a tokenizer. Tokenizers do not have prices; models and billing categories do.
The current API includes:
The current registry uses exact local token counts and dated direct OpenAI USD snapshots. The 2026-09-24 GPT-5 snapshot remains selectable after the 2026-09-25 expansion. The 2026-09-25 rates per million text tokens are:
| Model | Input | Cached input | Output | Source |
|---|---|---|---|---|
gpt-4.1 |
$2.00 | $0.50 | $8.00 | OpenAI model page |
gpt-4o |
$2.50 | $1.25 | $10.00 | OpenAI model page |
gpt-5 |
$1.25 | $0.125 | $10.00 | OpenAI model page |
The dates label the bundled observations, not provider-declared effective intervals. The API does not infer prices between snapshots. Before broadening the API further:
- [x] Define input, cached-input, and output categories explicitly.
- [x] Store price, currency, unit, provider, snapshot date, and registry version.
- [x] Pin pricing snapshots so results cannot change silently.
- [x] Allow an explicit caller-supplied pricing override for private or negotiated prices.
- [x] Resolve the model to its tokenizer independently from its price metadata.
- [x] Report the token-count mode (
exactorestimate) alongside the result throughestimate_cost_details(). Current supported models all useexact. - [x] Include the exact counted tokens in
estimate_cost_details()so the applied rate and cost can be audited per row. - [x] Keep raw-text cost estimates distinct from full request costs, which may also include wrappers, tools, images, audio, or provider serialization.
- [x] Test exact price-snapshot date boundaries and unknown/retired identifiers outside the pinned model registry. No historical price interval or live model-availability claim is inferred from the snapshots.
- [x] Avoid runtime network access in DataFrame expressions.
The initial operation should estimate the cost of the provided raw text in one declared billing category. Full request accounting remains a separate future feature.
Later DataFrame analytics¶
- Per-column and combined-row descriptions.
- Fused total/min/max/mean/histogram operations that avoid materializing a count column.
- Sampled total/distribution estimates with confidence intervals.