Skip to content

polars-llm

PyPI version Python versions Build status License

Call chat, TypeSafe decision, and embedding models from a Polars DataFrame, one row at a time, using native Polars expressions.

polars-llm registers an .llm namespace on Polars expressions so you can call any LangChain-supported chat model or embedding model on every row of a DataFrame — synchronously or asynchronously — and pipe the responses straight back into your data pipeline.

import polars as pl
import polars_llm  # noqa: F401  — registers the `.llm` namespace

(
    pl.DataFrame({"user_prompt": ["Summarise polars in one sentence."]})
      .with_columns(
          pl.col("user_prompt").llm.openai(model="gpt-4o-mini").alias("answer")
      )
)

Why polars-llm?

  • Expression-native — works inside with_columns, select, and any other Polars expression context.
  • Sync and async — aopenai, aanthropic, agemini fan out concurrently with asyncio.gather.
  • Provider-agnostic clients — chat / achat and embed / aembed accept any LangChain-compatible client.
  • Per-row prompts and system messages — every argument can be a Polars expression.
  • Structured outputs — pass a Pydantic schema as schema= and get a struct column back.
  • Typed decisions — run TypeSafe Choice, Score, and Noul questions together and get probabilities and confidence as nested structs.
  • Embeddings — openai_embed and gemini_embed return List[Float64] columns.
  • Vector search — compare vectors in an expression or run a top-K nearest-neighbor join between DataFrames.
  • Powered by LangChain.

Install

Python 3.10 or newer is required. Python 3.9 is no longer supported.

pip install "polars-llm[openai]"

Choose a different extra for Anthropic, Gemini, TypeSafe, or nearest-neighbor search. See Getting started for all installation options and environment variables.

Quickstart

Chat per row

import polars as pl
import polars_llm  # noqa: F401

df = (
    pl.DataFrame({"user_prompt": ["Capital of Spain?", "Capital of France?"]})
      .with_columns(
          pl.col("user_prompt").llm.openai(model="gpt-4o-mini").alias("answer")
      )
)

The result is an ordinary DataFrame with a new answer column. From there it can be filtered, joined, grouped, or written with the rest of your Polars pipeline.

Any LangChain-compatible provider

Use the generic verbs when a provider does not have a dedicated convenience method:

from langchain_ollama import ChatOllama

chat = ChatOllama(model="llama3.2")

df.with_columns(
    pl.col("user_prompt").llm.chat(client=chat).alias("answer")
)

Where next?

  • Getting started covers installation, authentication, and your first end-to-end pipeline.
  • Examples has recipes for prompts built from columns, structured extraction, concurrent calls, embeddings, and vector search.
  • API reference lists every expression and DataFrame method.