A new category of AI model is forming around a simple idea: not every problem needs text generation. Decision models — transformers trained to pick between defined options and assign calibrated confidence scores instead of producing prose — emerged into public awareness with TypeSafe AI’s proprietary Jev model in September 2026. Now AWS’s Strands Agents team has released Strands Decider 2B, the first significant open-source model in this category. It runs on a single consumer GPU, makes decisions in tens of milliseconds, costs nothing to operate after training, and is released under Apache 2.0 with full training data and code. For product teams building agentic AI workflows, this shifts decision models from a proprietary API to something you can run on your own hardware.

In Brief

  • Strands Agents released Strands Decider 2B on October 1, 2026, as an open-source decision model under Apache 2.0 license
  • The model is a 2-billion-parameter LoRA adapter on Qwen3.5-2B-Base, with a pointer head replacing the standard LM head — total added parameters are just over 1 million
  • It ranks 3rd of 33 models in the 2B class on JevBench public tasks, and 1st of 30 when excluding models just over the 2B threshold
  • Median latency is approximately 115ms on an NVIDIA RTX 3090 and 153ms on an M3 MacBook Pro
  • Training takes about 11 hours on a single RTX 3090 or roughly 1 hour 10 minutes on eight H100 GPUs
  • The release includes all training data, scripts, and evaluation code on GitHub, making it reproducible
  • Authors are Marc Brooker, Mike Chambers, and Fabio Nonato de Paula from the Strands Agents team

What Strands Decider 2B Actually Does

Like TypeSafe’s Jev, Strands Decider 2B does not generate text. It takes a state (text context) and a set of options, then returns a probability distribution over those options with a calibrated confidence score. The three primitives are the same as Jev’s: noul (yes/no questions), choice (selecting from a defined set), and score (rating on a scale).

The practical difference is that Strands Decider 2B is open-source, small enough to run locally, and comes with everything needed to retrain or modify it. You can install it via pip, run it on a consumer GPU or even a MacBook, and get answers in tens of milliseconds. No API calls, no per-token billing, no network latency.

The architecture is straightforward: take a pre-trained Qwen3.5-2B-Base model, remove the LM head (eliminating text generation capability), and replace it with a pointer head that scores the hidden states at each option position against the hidden state at the answer position. The pointer head is small — just over a million parameters. The base model is fine-tuned with a rank-16 LoRA adapter, keeping the training footprint manageable. The released version is v19, meaning the team iterated through 18 previous architectures to get there.

According to the Strands Agents blog post, the current version evolved from an earlier architecture that used a “slot head” instead of a pointer head. The team found the slot head performed significantly worse and switched. This kind of iterative architecture search is documented in the repository, which makes the release valuable as an educational resource even for teams that never use the model directly.

How It Performs

The Strands team reports three metrics that matter for decision models: accuracy, calibration, and latency.

On JevBench’s public set (231 tasks), Strands Decider 2B achieves 72.29% accuracy (167 of 231 correct). The team states this places it 3rd of 33 models in the 2B parameter class, and 1st of 30 when excluding models that are just over the 2B threshold. The model also achieves 100% accuracy on JevBench’s easy tasks — the ones that map to routine decisions in agentic workflows.

On internal evaluation sets, the results vary by task type. ContractNLI (contract non-disclosure agreement reasoning) achieves 87.2% accuracy across 1,026 tasks. MuSiQue (multi-step question answering) reaches 88.4% across 1,199 tasks. Boardgame-QA scores 82.2% on 900 tasks. HotpotQA (held out) scores 71.7% on 959 tasks. These are not frontier-LLM numbers, but they are not meant to be — they are numbers for a 2B model making typed decisions in milliseconds.

For latency, the model returns decisions in a median of approximately 115ms on an NVIDIA RTX 3090, with latency increasing approximately linearly with task size (measured in tokens). On an Apple M3 MacBook Pro, the median is around 153ms for small tasks. This is fast enough to sit inside real-time agent workflows where an LLM call would introduce unacceptable delay.

Why This Matters for Product Teams

The TypeSafe Jev article on this site covered why decision models matter for product architecture: they handle the invisible decision points in AI workflows — content moderation, intent classification, safety checks, routing, escalation gates — at a fraction of LLM cost. Strands Decider 2B changes the calculation in three ways.

No API dependency. Jev is a hosted API. Strands Decider 2B runs on your own hardware. For teams in markets where API costs are a significant barrier — including Pakistani technology teams paying dollar-denominated LLM prices — a model that runs locally on hardware they already own eliminates a category of cost entirely. There is no per-call billing. There is no vendor lock-in. There is no data leaving your infrastructure.

Reproducibility and customization. The release includes all training data, training scripts, and evaluation code on GitHub. A team that needs a decision model for a specific domain — say, classifying student queries on an EdTech platform, or routing customer support tickets for a Pakistani fintech — can retrain or fine-tune Strands Decider 2B on their own data. The training cost is modest: about 11 hours on a single RTX 3090, which is the kind of GPU available on the used market for under $1,000.

The hybrid agent pattern. The Strands team describes a hybrid approach where an LLM handles the hardest decisions and a decision model handles the routine ones. In their demo, an agent that eagerly calls a weather tool without knowing the user’s city is intercepted by Strands Decider 2B, which checks two yes/no questions before the tool executes: are the argument values grounded in what the user actually said, and is it premature to call this tool? If the decision model says no, the agent goes back to ask for clarification instead of making up a location. This pattern — cheap decisions gating expensive LLM calls — is where decision models deliver the most value.

The Emerging Decision Model Ecosystem

Strands Decider 2B does not exist in isolation. The decision model category is forming rapidly:

  • TypeSafe Jev (September 2026) — proprietary, hosted API, highest ranked on JevBench v1.3.0 with a composite score of 74.4
  • Strands Decider 2B (October 2026) — open-source, local execution, 3rd in the 2B class on JevBench
  • Open-weight alternatives based on Qwen, Gemma, and other foundation models are appearing in JevBench rankings, suggesting the pattern is replicable
  • OpenAI’s Decisions API provides similar primitives (check conditions, select from options, score against rubrics) integrated into the OpenAI platform

The fact that multiple organizations are converging on the same architecture — transformer torso, no LM head, calibrated output over defined options — suggests this is not a one-off experiment. It is a category forming around a genuine gap in the AI model stack.

For context, the hidden costs of AI coding agents are not just about generation tokens. They are also about the compounding cost of using expensive models for simple decisions. Every routing check, every safety gate, every “should this agent proceed?” question that runs through a frontier LLM is a decision that a 2B model could handle at a fraction of the cost — or, with Strands Decider 2B, at no cost at all beyond the electricity to run a GPU.

What About Pakistan and Emerging Markets?

The open-source, locally-runnable nature of Strands Decider 2B matters more in markets where API costs are a structural barrier. For a Pakistani EdTech platform that needs to classify thousands of student interactions per day — routing queries, checking whether a question is appropriate, determining whether a student needs human support — the choice between a frontier LLM at $5-15 per million tokens and a local 2B model that costs nothing after the initial hardware investment is not marginal. It is the difference between a feature that is economically viable and one that is not.

This connects to a pattern visible across technology adoption in developing markets: the same capability becomes transformative when the cost drops by an order of magnitude. A locally-runnable decision model does not just reduce costs — it changes what is buildable. Teams that could not afford to run real-time safety checks on every student interaction can now do so. Teams that could not justify the API spend for agent routing can now build hybrid agent workflows without a per-call budget.

The training cost matters here too. A Pakistani university lab or startup with a single RTX 3090 can retrain Strands Decider 2B on local-language data or domain-specific tasks in less than a day. That is a materially different proposition from fine-tuning a frontier LLM, which typically requires cloud compute costing hundreds or thousands of dollars per run.

Limitations and What the Data Does Not Tell You

The Strands team is transparent about limitations, which is worth noting because it is not always the case with AI releases.

Long multi-step documents are the weak spot. The team explicitly states that JevBench’s hard tier scores far below the easy tier, and that questions are read less than documents — meaning that with the state and options fixed, a changed question often gets the same answer. This suggests the model is better at reading comprehension than at question understanding, which limits its use in tasks where the phrasing of the question is the key variable.

Calibration is one temperature per primitive. The confidence bands are established on classification tasks specifically. The team recommends measuring calibration on your own traffic before trusting a threshold. In practice, this means you cannot assume that a 0.8 confidence score means the same thing on your data as it does on JevBench — you need to validate.

Training data inherits label noise. The model is trained on a mix of open datasets (arXiv classification, CLINC intent detection, Banking77, BoolQ, emotion classification, and many others) plus synthetic data generated by open-weight language models. The team notes that the training mix includes domains and label noise from these sources. For production use, teams should evaluate on their own data rather than assuming benchmark performance transfers.

It cannot reason. Like Jev, Strands Decider 2B is a System One model — fast, pattern-matching, single-pass. It cannot do chain-of-thought reasoning, cannot break down complex problems, and cannot generate explanations. For decisions that require multi-step reasoning, an LLM is still the right tool.

Product Builder’s Perspective

From a product-building perspective, Strands Decider 2B represents something specific: the moment when a new architectural pattern becomes accessible to teams without deep ML research budgets.

The architecture — transformer torso with the LM head replaced by a pointer head — is not conceptually difficult. But implementing it well requires iterative architecture search, training pipeline engineering, calibration measurement, and evaluation infrastructure. The Strands team went through 19 versions to get to this point. By releasing the code, data, and version history, they have made it possible for other teams to start from v19 instead of v1.

For teams building educational technology platforms or AI-powered products in emerging markets, the practical question is not whether to use Strands Decider 2B specifically. It is whether the decision model pattern belongs in their architecture. The answer depends on how many classification, routing, and gating decisions their product makes per user interaction. If the answer is “dozens” — which it is for any product with agent workflows, content moderation, or automated support — then a decision model that runs locally at no marginal cost is worth evaluating.

The comparison to the early days of embedding models is apt. Before open-source embedding models like sentence-transformers became widely available, teams used LLMs for semantic search and similarity tasks because that was the only option. Once purpose-built embedding models became accessible, the industry moved to a cheaper, more appropriate tool for that specific job. Decision models are the same pattern, one layer up: a purpose-built tool for the decision layer that sits underneath agentic workflows.

What to Watch Next

  • JevBench rankings. The benchmark is evolving and more models are appearing. Watch whether open-source decision models close the gap with proprietary offerings like Jev over the coming months.
  • Domain-specific fine-tunes. The training infrastructure is open. Expect community fine-tunes for specific verticals — customer support, education, healthcare — to appear on Hugging Face.
  • Integration with agent frameworks. The Strands team is working on libraries for decision model integration. Similar patterns will likely appear in LangChain, CrewAI, and other agent frameworks as decision models gain adoption.
  • Hardware requirements. A 2B model that runs on a MacBook is already accessible. If quantized versions (INT4, INT8) appear, the hardware barrier drops further — potentially bringing decision models to Raspberry Pi-class devices, similar to how 1.58-bit BitNet quantization brought LLM inference to ESP32 microcontrollers.

For product teams, the recommendation is simple: if you are building agentic AI workflows and using LLMs for routing, classification, or safety checks, evaluate Strands Decider 2B against your current approach. The cost of evaluation is a few hours of engineering time. The potential savings — in API costs, latency, and architectural simplicity — are ongoing.

Sources