AWS Strands Decider 2B: Open Decision Model for Agents

AWS Strands Decider 2B: Open Decision Model for Agents

Not every step in an AI agent needs a model that writes. Many steps only need a model that picks. That is the idea behind Strands Decider 2B, a new open-source release from Amazon Web Services' Strands Labs team. AWS describes it as a lightweight "decision model" built to make fast choices inside agentic workflows without generating any text.

The model is out now on Hugging Face. The full codebase, training scripts and examples are on GitHub. AWS has tuned it for local use and quick experiments, so developers can run it on their own laptops as well as in public cloud environments.

What a decision model actually does

Decision models, sometimes called "System 1 models," work differently from the large language models most people now know. An LLM takes an input and produces something new, such as text, code, images or video. A decision model is given a set of predefined options and returns one of them. That is its only output.

Each choice comes with a confidence score. The score matters more than it might seem. Because the model produces no text, it cannot explain why it chose what it did. The score is the only signal users get about how reliable a given answer is likely to be.

Dropping text generation has a clear benefit: speed. Generating text uses a lot of tokens and adds latency. A model that only selects from a list avoids both costs.

Why now

AWS says decision models are not new, but they drew little attention until recently. That changed a couple of weeks ago, when a startup called TypeSafe AI Inc. released Jev. Jev was designed to make fast, structured decisions that software and AI agents can use directly.

AWS credits Jev with showing how useful this class of model can be. It also says Jev has structural weaknesses. The main one, in AWS's view, is a parallel output structure that causes performance problems when the model has to handle intricate reasoning. Strands Decider 2B is the company's first serious attempt to improve on the concept.

How AWS built it

The model starts from a standard LLM. AWS took the base torso of Qwen3.5-2B and removed the usual "LLM head," the part that normally produces text. In its place sits a small custom "pointer head" with roughly 1 million parameters.

The team then fine-tuned the Qwen3.5-2B torso with a rank-16 LoRA adapter. LoRA is a common technique that adapts a model by training a small set of added weights rather than the whole network. Here, it teaches the model to score the hidden states of the available choices directly against answer positions.

AWS says it went through many iterations before this release. The 2-billion-parameter size was a deliberate choice. According to the company, it is small enough to run on a local machine with latency under 150 milliseconds, but still capable enough to handle complex decisions. If you have been following smaller models running on local hardware, the Qwen base will be familiar.

On benchmarks, AWS reports that Strands Decider 2B scored highly on accuracy and calibration on JevBench, compared with other open-source models of the same 2B size. Calibration measures how well a model's confidence matches how often it is actually right. For a model whose confidence score is its only form of explanation, that is the number to watch.

Where it fits in an agent

AWS lists several tasks where it expects the model to help:

  • Model routing: sending a request to the right model.
  • Tool selection: choosing which tool an agent should call.
  • Context management: deciding what information to keep or pass along.
  • Guardrail enforcement: checking whether an action should be allowed.
  • Policy classification: sorting inputs against defined rules.

The company also sees the release as a step toward "hybrid agents." In that setup, a decision model handles simple, repetitive choices, and an LLM takes over when a problem needs real reasoning. AWS says its broader goal is to speed up agentic AI development across the developer community.

Our Take

This release suggests that agent builders are starting to break their systems into specialized parts, not route every step through one large model. Many agent steps are really multiple-choice questions: which tool, which model, allow or block. Using a full LLM for those is slow and expensive. A small model that answers in under 150 milliseconds could change the economics, which fits a wider push toward cheaper models that come close to flagship performance.

The guardrail use case deserves a closer look. Placing safety checks for agents in a separate, fast component is appealing. But a decision model cannot explain itself, so teams will rely heavily on how well its confidence scores are calibrated.

Several things are worth watching. Will independent tests confirm AWS's JevBench results? How will TypeSafe AI respond with future versions of Jev? And will hybrid agents become a common pattern, or stay a niche optimization? The open release means developers can find out for themselves.