Skip to content

Exact token counts in Polars

polars-tokenizer adds native, count-only token expressions to Polars. It counts raw UTF-8 text without building token-ID arrays and returns a UInt32 column that can be used in eager, lazy, and streaming queries.

import polars as pl
import polars_tokenizer as tokens

frame = pl.DataFrame({"text": ["hello world", None, ""]})
counts = frame.select(tokens.count("text").alias("token_count"))
assert counts["token_count"].to_list() == [2, None, 0]

The default tokenizer is o200k_base. You can instead choose cl100k_base, p50k_base, or r50k_base, or pass a supported model alias. String, UTF-8 Binary, Categorical, and Enum columns are supported. Nulls stay null, and special-token-looking text is counted as ordinary text.

Start with the getting-started guide, then see the API reference for count, cost, cache, and pricing options. The Binary input guide explains UTF-8 validation and nulls.

For implementation and reproducibility details, see the architecture, benchmarking protocol, and provider capability matrix.

The library counts only the supplied raw text. It does not estimate a complete provider request, including roles, wrappers, tools, images, or audio.