Decision models¶
Note
Experimental: this can change or be removed in a minor release of this provider. See Stable and experimental features.
Every other model in this provider writes text. A decision model does not: you give it some text and a typed question, and it answers with a value from a set you named in advance, plus a confidence. Ask it for a string and the request is refused before it leaves your process.
pydantic-ai reaches decision models through two model prefixes. Nothing in this provider is
specific to either: both arrive through the same
PydanticAIHook as every other
model, so a model id is the whole integration.
Prefix |
Needs |
Runs |
|---|---|---|
|
|
TypeSafe’s hosted Jev. |
|
|
Any server that answers the same |
Both prefixes need a newer pydantic-ai-slim than the floor the provider’s other extras
set, so check the release you have installed.
Setup¶
The examples on this page use a connection named decision_default. Create it for the
backend you run.
TypeSafe Jev
Install the extra, which adds the TypeSafe SDK:
pip install 'apache-airflow-providers-common-ai[typesafe]'
Create a connection (
Admin > Connections):Connection Id:
decision_defaultConnection Type:
Pydantic AIPassword: your TypeSafe API key, from your TypeSafe account
Extra:
{"model": "typesafe:jev-1.13.0"}
Leave Host empty unless you are pointing at a proxy; the provider defaults to TypeSafe’s own endpoint.
A System One server
Start the server, or note the URL of a hosted one. For example, to serve Strands Decider on your own machine:
pip install strands-decider strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000
Create a connection (
Admin > Connections):Connection Id:
decision_defaultConnection Type:
Pydantic AIHost: the server’s URL, such as
http://decider.internal:8000, with or without a trailing/v1Password: the server’s API key, sent as a bearer token. Leave it empty for a server that takes none, such as a local Ollama.
Extra:
{"model": "system-one:strands-decider-2B-hobson-v19"}
The name after
system-one:is sent to the server as the model to answer with. A server that runs several, such as Ollama, picks one by it.Describe every option you ask about, and give the question its text. Servers differ in what they accept, and Strands Decider 0.1.0 refuses, with an HTTP 422 that fails the task, a question that has an option without a description or that has no question text at all. Describe each branch in
branchesand each member of anEnumthrough a docstring; the members of a bareLiteralhave no descriptions, so use a describedEnuminstead, as the worked example below does. The question text is the operator’ssystem_promptor the agent’sinstructions, which default to empty.
Pin the model and its version, as in typesafe:jev-1.13.0 rather than
typesafe:jev-latest. A threshold you tuned against one model, or one release of it, is
not guaranteed to mean the same thing on another, so measure it again when you change
either.
When a decision model is the right choice¶
All four of these have to hold.
The answer is a label, a number, or a pick, not prose. A bool, a Literal or
Enum of strings, a bounded float, a whole-number rubric, or a list of picks. A
str field is refused, so anything that writes a summary, a query, a migration, or a
reply to a person is out.
You can name the options up front, and there are few enough for the model. Downstream
task ids, error categories, severity levels, environments, a repository list. If the set is
open, or is discovered at run time and could grow past the cap, this is the wrong tool. Each
model has its own cap per question: Jev takes 255 options, and Ollama takes 26. Tools
attached to an agent count as options in a question of their own, with the output type as
one more. pydantic-ai refuses a question over Jev’s cap before sending it; a system-one:
model’s cap is the server’s, so a question over it comes back as an HTTP error from the
server instead.
The decision is on a path where latency or cost is the constraint. One decision per Dag run rarely justifies changing models. One per row, per file, or per retrieved document does, and so does a branch a scheduler is waiting on.
You want a number to branch on, not a sentence to trust. A pick, a rubric, and a
yes/no each come back with a confidence, so “act automatically above 0.8, ask a human below
it” becomes something you can write down. A bounded float is the exception: there the
probability is the answer, so nothing separate is reported and you gate on the value
itself. If you would not do anything different with a confidence of 0.55 than with 0.95,
that is a sign a general-purpose model is fine here.
And one case where the answer is neither: if a deterministic rule already sorts the input correctly, use the rule. A decision model is cheap, not free, and a rule you can read is worth more than a probability you have to calibrate.
Where it fits in this provider¶
Surface |
Fits |
Why |
|---|---|---|
Yes, with a caveat |
The downstream task ids are already presented to the model as a constrained set of
choices, which is exactly the shape a decision model answers. Pointing
|
|
|
Yes |
A |
Yes |
The model names one of the policy’s |
|
Agents with toolsets |
Partly |
Which tool the text calls for is itself a pick, so a decision model can make it.
What it cannot write is a tool’s arguments. A tool taking none it calls itself; one
taking arguments raises |
Reading the confidence¶
The confidence lives in provider_details on the model response. It is reported per
output field, and a bare output type is reported under "response". A bounded float
field reports none at all, because there the probability is the answer.
Two surfaces act on it for you. LLMBranchOperator
and LLMOperator take a
decision_policy whose min_confidence sends an unsure answer to a person, or fails
the task, before anything downstream runs on it, and record the confidence, the
probabilities and the bar in the decision XCom (see Branch on an answer: LLMBranchOperator and @task.llm_branch).
ClassifierRetryPolicy takes the same min_confidence and hands an
unsure answer to fallback_policy, then its deterministic fallback rules. In the branch operator and the retry policy, a
per-option bar lets the choice whose wrong pick costs most demand more certainty than the rest.
Outside those, read it yourself. AgentOperator carries it inside the message_history
transcript when that is enabled, and a hook-level call has it on the result:
from enum import Enum
from pydantic_ai import UseEnumMemberDocstrings
from airflow.providers.common.ai.hooks.pydantic_ai import PydanticAIHook
from airflow.sdk import task
class FailureCause(UseEnumMemberDocstrings, str, Enum):
transient = "transient"
"""A fault that clears by itself: a timeout, throttling, a dropped connection."""
resource = "resource"
"""A dependency is down or unreachable and needs fixing before a retry can work."""
permanent = "permanent"
"""A bug or bad input that fails the same way however often it is retried."""
@task
def triage(log_line: str) -> dict:
agent = PydanticAIHook(llm_conn_id="decision_default").create_agent(
output_type=FailureCause,
instructions="Classify why this Airflow task failed.",
)
result = agent.run_sync(log_line)
details = result.response.provider_details or {}
confidence = (details.get("confidence") or {}).get("response")
return {"category": result.output.value, "confidence": confidence}
Then branch on the returned confidence in a downstream task, so an unsure answer escalates instead of acting. Use a higher bar for acting automatically than for flagging something for review: collapsing both into one number is the easier thing to tune and the wrong shape for the decision.
Worked example¶
example_decision_model.py has both halves: a branch with nothing specific to decision models but its connection, and a classification that escalates when the confidence is low.
What it answers badly¶
Read pydantic-ai’s decision model guide and the documentation of the model you run (TypeSafe’s for Jev) before you trust a number from one of these models. Two failure modes documented for Jev matter more than the rest in a Dag; check them against any other model before relying on it not to share them:
The text is treated as data, not as hostile. An injected instruction, a misleading framing, or an argument for its own answer can move the result. A guard built on a decision model belongs alongside deterministic checks, not instead of them.
Option order is part of what the model sees. Reordering a
Literal’s members can change the answer, so a threshold measured against one ordering is measured against that ordering only.