Every tool call your agent makes today costs a full LLM round trip. Strands Decider 2B is a 2B-parameter model built to do the cheap ones instead, and it runs on the laptop you already have.

What It Actually Is

Strands Decider 2B takes a pretrained Qwen3.5-2B torso and removes the LM head — the part that generates text. It replaces it with a pointer head of just over a million parameters that scores each offered answer by comparing the hidden state at each option position against the hidden state at the <answer> position. The torso is fine-tuned with a rank-16 LoRA adapter.

That means it cannot write prose. It picks between options and assigns confidence scores. Per the announcement: "decision models are designed to pick between sets of options" and "assign simple numerical scores." In exchange you get speed, a reliability score per decision that frontier inference APIs do not expose, and the ability to ask many questions about the same prompt efficiently.

The tradeoff is real. The single-parallel-pass design makes it "significantly worse at solving complex problems than reasoning models," and the lack of text generation rules out coding, chatbots, and summarization.

The Numbers

Metric Value Hardware
Parameters 2B (torso + ~1M-param head) —
Median decision latency ~115 ms RTX 3090
Median latency, small tasks ~153 ms M3 MacBook
JevBench accuracy rank 3rd of 33 in 2B class —
Ranking excluding just-over-2B models 1st of 30 —

Latency scales approximately linearly with task size in tokens. The published latency figures are against v18 of the model; the release itself is v19. Everything changed between versions is documented in the repo.

I would not run a SolverBench-style task on this. The claim that matters is 100% of the easy JevBench tasks answered correctly — that is the rote-decision territory where I would actually deploy it.

Where It Fits In An Agent

The repo ships an example inside a Strands agent. The agent runs locally, connects to the decider also running locally, and uses the default LLM from Amazon Bedrock for the hard parts.

The scenario is deliberately small. The agent has a get_weather tool and an eager system prompt, so when you ask "What's the weather?" without a city it guesses one and calls anyway. Before the call runs, the decider answers two yes/no questions: are the argument values grounded in what the user said, and is it premature to call this tool. The policy is wired through the InterventionHandler with a before_tool_call method, passed to Agent(interventions=[...]), returning a typed action: Proceed, Deny, Confirm, or Guide.

A few lines of Python convert predictions into that action, and the agent asks which city you meant instead of inventing one. The pattern is the point: a decision this cheap can sit in a path where an LLM call never could.

Note that the questions, threshold, and policy in the example were all picked by hand. Treat it as an illustration, not a recommendation.

Reference: Getting It Running

Install:

pip install strands-decider

Ask a choice question with state:

strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days! " \
  --choice "Which team should handle this?=billing,sales,retail"

Example output:

choice_0 -> billing (confidence 0.768)
billing                  0.845
retail                   0.091
sales                    0.064

Weights are on Hugging Face, and the GitHub repo includes all training data and scripts plus agent examples under examples/strands/. The Strands team says it is working on libraries for decision model integration, so expect the API shape to move.

What I Would Watch

The class is new — TypeSafe AI's Jev launched earlier this month — and the tooling around it is thinner than the model. The 115 ms median is a floor that the team itself says it wants to lower. If you are already gating tool calls with hand-written heuristics, this is a drop-in upgrade worth an afternoon. If your agent's bottleneck is reasoning quality rather than call gating, this changes nothing for you yet.

Sources: