A 1.9B "decision model" that answers faster than an LLM and tells you how sure it is
What this project does
In one sentence: it's a brand-new class of "decision model" (also called a "system one" model) built for the jobs that LLMs do poorly and traditional classifiers make too heavy — picking an option from a set, or scoring something on a scale, every time with a calibrated confidence score. Feed it a sentence like "My payouts have been failing for three days!" and ask "Which team should handle this? = billing, sales, retail" and it replies "billing, confidence 0.768". Ask "Does this convey urgency?" and it returns a number between zero and one. Ask "How frustrated is the writer? = calm, frustrated, depressed" and it gives you a scored scale. Essentially it is a lightweight engine for the "judgement" steps inside an agent workflow, pulling those repeating micro-decisions out of expensive large-model calls.
Why it's blowing up
The repo landed on 2026-09-29 and reached 294 stars within days. The logic is clear: agents are everywhere now, but inside an agent a huge amount of micro-decisions — "choose A or B", "which tool should I call", "what's the sentiment of this sentence" — are expensive and slow if you call a full-size LLM, yet training your own classifier takes data and effort. strands-decider sits in the middle: it responds far faster than an LLM (a median of 115 ms on an RTX 3090) and needs no dedicated training like a classic classifier. The killer feature is the confidence score: on short classification tasks it has never seen, when it says confidence 0.9 or above, it is right about 95% of the time, and below that it tells you to confirm or ask a human. That "knows-when-it's-unsure" ability is simply not exposed by frontier LLM inference APIs, yet it is exactly what production systems need — you always want to know when to hand a decision back to a person, and a raw LLM logprob is not the same thing as a calibrated probability you can trust.
Technical highlights
At its core is a 1.9-billion-parameter model (strands-decider-2B) that answers in a median 115 ms and runs on an RTX 3090, an Apple-silicon Mac, or even plain CPU. Usage is straightforward: the CLI strands-decider ask
Who it's for
Engineers building agent frameworks, routing and tool-selection will feel this most. Imagine letting the agent itself decide "should I call search or write code", "should this user message be escalated to a human or auto-replied", "is this review positive or negative" — these high-frequency micro-decisions cost an order of magnitude less with this than calling an LLM. ML researchers can also use it as a ready-made sample of the emerging "system one model" direction, skipping the trouble of training a classifier from scratch. It runs on a single machine and serves even without a GPU, so deployment is light and it slots neatly into an existing pipeline as a judgement middleware.
Quick start
Install with pip install strands-decider. Run a decision: strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 --state "Help! My payouts have been failing for 3 days!" --choice "Which team should handle this?=billing,sales,retail". To serve it: strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000. For images, pip install "strands-decider[vision]" then add --vision. The model weights themselves are pulled from the StrandsAgents repo and downloaded automatically on first run.
How it compares
Against a full LLM answering directly: an LLM can generate arbitrary text but is slow, costly, and gives no calibrated confidence; the decider does only "choose/score" but dozens of times faster and tells you how sure it is. Against a traditional sklearn classifier: the classifier needs labeled data, feature engineering and training, while the decider is general-purpose and pre-trained, ready to use, filling the gap "between LLMs and classifiers". It is a sibling of the Strands Agents SDK but the model itself is generic and any agent framework can plug it in. If an LLM is "system two" (slow thinking), this is "system one" (fast intuition), built to catch the millisecond-level decisions inside an agent and let the big model focus on the deep reasoning it is actually good at. In a real pipeline the two complement each other: the decider triages and routes in real time, and only the hard cases ever reach an expensive LLM call, which is exactly how you keep an agent both fast and cheap.