Getting started¶
Install¶
Install the prebuilt wheel for a supported platform and CPython 3.10–3.14:
The package pins the compatible Python Polars version. If no wheel exists for
your platform, pip will try to build the Rust source distribution instead.
Install from this repository¶
The project is currently built as a native Python extension. From a clone,
install Rust 1.87+ and uv, then run:
The Python and Rust Polars dependencies are pinned together because native expression plugins use Polars' plugin ABI.
Count text¶
Importing polars_tokenizer also registers the Expr.tokens namespace. The
functional API works with a column name or Polars expression:
import polars as pl
import polars_tokenizer as tokens
frame = pl.DataFrame({"text": ["hello world", None, ""]})
counts = frame.select(tokens.count("text", tokenizer="o200k_base"))
assert counts["text"].to_list() == [2, None, 0]
The namespace API composes naturally with lazy and streaming plans:
result = (
frame.lazy()
.with_columns(pl.col("text").tokens.count(model="gpt-5").alias("token_count"))
.collect(engine="streaming")
)
Use either tokenizer or model, not both. The default tokenizer is
o200k_base. Model aliases are exact, versioned mappings rather than prefix
matches; see provider support.
Count repeated values¶
Categorical and Enum columns automatically count each used dictionary value once. For String or Binary columns with many repeated values, an optional bounded FIFO cache can improve throughput:
The capacity is 1–65,536 entries per worker task. Leave caching off for mostly unique text, where hashing and eviction can be slower.
Estimate raw-text cost¶
For a supported model, cost expressions multiply exact local counts by a
pinned USD-per-million-token price. The category is input by default; the
other choices are cached_input and output.
costs = frame.select(tokens.estimate_cost("text", model="gpt-5", category="input"))
details = frame.select(tokens.estimate_cost_details("text", model="gpt-5"))
estimate_cost() returns a Float64 USD expression. The structured form also
includes the exact token count, count mode, price, snapshot date, registry
version, and source. Use price_info() to inspect the
bundled rate before building an expression.
These are raw-text estimates, not provider invoices or request-level accounting. They exclude messages, roles, tools, images, audio, and other provider serialization.