Automation and Agents
Strands Decider 2B: Which Agent Decisions Are Worth Handling Locally?
Strands Decider 2B selects supplied options locally. Use five questions and three worked cases to decide which agent choice merits a contained trial.

On this page
- One request can contain several decisions
- What Decider 2B supplies
- Screen a candidate decision with five questions
- 1. Can you name the answers and their actions?
- 2. Does the decision see the information it needs?
- 3. Can reviewers establish a reference answer?
- 4. What does each direction of error change?
- 5. Who handles an uncertain case?
- Apply the screen to three choices
- Route a routine request
- Check a proposed tool call
- Resolve an ambiguous payout problem
- Read the published results within their limits
- Run one contained trial
A customer writes, “My payouts have failed for three days.” An agent may need to choose a support queue, ask for missing information, and check a proposed tool call. Those are separate decisions with different consequences. A wrong queue may delay help. A guessed account identifier could send the agent toward the wrong record.
Strands’ announcement, dated October 1, 2026, introduces Decider 2B as a model that selects from supplied options or scores a supplied scale. It does not generate the customer reply. Its useful role is a bounded choice inside a larger workflow, where people define what each answer permits.
Start with one recurring decision whose answers are explicit, examples can be reviewed, mistakes have understood consequences, and uncertain cases have somewhere to go. A routine route may qualify for a local trial. An account change does not become safe merely because the model assigns a high score.
One request can contain several decisions
Consider a hypothetical agency that handles billing and account questions for clients. A customer reports failed payouts. The agency might route the message to billing, account support, or an ownership-review queue. It might then decide whether it has enough information to investigate. If an agent proposes a tool call, the agency might check whether its arguments come from facts the workflow is authorized to use.
The answers need not agree. Billing may own the case while the account identifier remains unknown. A three-day failure may affect priority without identifying a transaction. A number in an earlier message might refer to another account or appear inside a quoted example. A correct route does not validate that number.
Write each choice at the point where its answer changes the work. Record what information is available, which answers are allowed, and what action follows each answer. A routing label can assign an investigation owner. A missing-information decision can prompt a specific question. A tool-argument check can permit, stop, or escalate a call under the agency’s own rules.
For the hypothetical router, the decision record could read: “Given the customer’s latest request and the written queue rules, choose billing, account support, or ownership review. The answer assigns an investigation owner. It does not authorize an account change.” That final sentence keeps a narrow routing answer from acquiring authority over a later action.
View image detailThis separation also tells reviewers where to look when something goes wrong. A misrouted message points to the ownership rule or the text available to the router. An unsupported tool argument that passes points to the grounding rule, the proposed call, or the evidence shown to the check. One overall label of “handled correctly” would hide that diagnosis.
The exercise may expose a process problem before anyone runs a model. If billing and account support disagree about requests that combine failed payouts with lost access, the agency needs a shared-ownership rule. A model can select a label even when the organization has not settled what that label means.
What Decider 2B supplies
The Strands repository describes three question types. A choice selects one of the options supplied with the request. A yes/no question assesses a stated condition. An ordered score uses levels in a supplied rubric. Several questions can be asked about the same input.
These formats serve different jobs. A choice could propose a support queue. A yes/no question could flag a proposed call for clarification. An ordered rubric could describe urgency if reviewers agree on what each level means. The output format does not define “urgent” for an agency or decide who is allowed to act on that label.
Strands says it replaced a language model’s text-generation head with a component that scores offered answers. “2B” is the product name; the repository specifies 1.9 billion parameters. The practical distinction is the resulting answer: Decider selects or scores supplied options. Strands’ announcement says it is unsuitable for chat, coding, document summarization, and other work that requires generated text.
View image detailAn existing agent could still write a reply through its language-model path. Decider could answer a separate question just before a tool call executes. The workflow would then apply its own policy to that answer. The model supplies a decision output; it does not supply the organization’s policy.
In pages captured on October 3, 2026, the code repository was marked public and the v19 model card described a checkpoint. The repository states Apache-2.0 licensing for the project, and the card states the same for the checkpoint. The card says separate Qwen base weights download on first use and documents CPU, CUDA, and MPS device choices. The base-weight terms, current download access, installation, and execution were not checked for this article.
Screen a candidate decision with five questions
This screen is a proposed way to choose a useful trial. It is not a measured capability of Decider 2B. Each question identifies something the agency must define or observe before an automatic route can be judged.
1. Can you name the answers and their actions?
“Send this to the right person” leaves the destination undefined. “Choose billing, account support, or ownership review” gives reviewers three outcomes they can inspect. Each needs an action. Billing may receive the case immediately. Ownership review may ask for a missing detail. Someone must update the choices when team responsibilities change.
An “unclear” option can represent a real result, but adding it does not prove the model will use it well. If many cases span two teams, the agency may need a shared-ownership rule. If responsibilities shift every month, maintaining the labels, examples, and downstream routes belongs in the trial’s operating cost.
2. Does the decision see the information it needs?
A short payout message may support a tentative route. It may not identify an account or justify a transaction change. Specify the input at the decision point: the latest message, relevant conversation turns, a proposed call, or an authorized account record. If the correct answer depends on information outside that input, tidy labels will not fix the question.
Longer input can create its own ambiguity. A conversation may include corrections, quoted text, two accounts, and several requests. The agency may need to split the case or obtain a verified fact before asking for one answer. The v19 model card identifies long, multi-step documents as a weak spot, which makes that distinction especially relevant to a trial.
View image detail3. Can reviewers establish a reference answer?
Collect examples from the traffic the decision would actually receive. Give responsible reviewers a written rule and enough context to apply it. Record disagreements. If reviewers disagree about whether a failed payout with lost access belongs to billing or account support, forcing one label into the test set will make an accuracy figure look more settled than the underlying process.
Include routine and difficult cases. Routing examples might combine billing and access problems. Tool-call examples might contain an account number corrected later, a number quoted from someone else, or a fact supplied through a separate authorized step. Reviewers may reasonably mark some cases for clarification or ownership review. Keep those outcomes in the evaluation rather than dropping the inconvenient examples.
The reference process needs its own quality check. If reviewers routinely lack the information needed to label a case, record that gap. It may show that the workflow needs better intake or a different decision point. The Rise guide to choosing what to automate applies a similar test to repeated work: examine input stability, judgment, failure consequences, and maintenance before handing a task to a system.
4. What does each direction of error change?
A wrong queue and an unsupported tool argument have different consequences. For each possible answer, describe what happens when it is granted incorrectly and when it is withheld incorrectly. Who notices? Can the action be reversed? Does it add a handoff, delay a response, expose information, or alter a record?
Suppose a router handles common billing messages well but sends unusual, time-sensitive cases to a slow queue. Overall accuracy could hide the cases the agency most needs to catch. A tool check that blocks almost every call may avoid unsupported calls while creating an unmanageable review queue. Count error types and their consequences before deciding whether the average is useful.
View image detail5. Who handles an uncertain case?
Define the fallback before allowing automatic action. The workflow could ask the customer a targeted question, use the existing decision path, or send the case to a monitored review queue. Name the owner of that destination and the facts they need to resolve the case. “Escalate” is incomplete when no one owns the escalated work.
A confidence threshold might help select fallback cases after it has been evaluated on representative traffic. The model card says its calibration was fitted on held-out short classification tasks and that the established confidence bands apply there. It advises measuring calibration on intended traffic before trusting a threshold. The agency must find out whether scores distinguish reliable from unreliable answers on its own messages.
Fallback volume matters as much as the threshold itself. If a proposed setting sends half the queue to a reviewer, the reviewer’s time becomes part of the new route. If the setting lets costly mistakes through, a low fallback count is no victory. The team needs both numbers, plus the consequences of the cases in each group.
If an agency cannot define the answers, provide the needed facts, review representative examples, describe errors, or staff the fallback, it has not yet specified an automatic decision. Clarifying those parts may improve the workflow even if it never uses Decider 2B.
Apply the screen to three choices
Route a routine request
Imagine the agency has written ownership rules. Billing handles failed payouts. Account support handles access problems. Ownership review handles messages whose destination cannot be determined from the available information. A local decision would select one of those three destinations. It would assign an owner, not grant permission to inspect or alter an account.
The agency could label past messages under those rules, then compare proposed routes with reviewed destinations. A false billing route could delay an access fix. A false account-support route could delay a payout investigation. Sending clear cases to ownership review would consume staff time; sending unclear ones directly to a specialist could create repeated handoffs.
This is a plausible trial because the options, input, and likely consequences can be described. It still needs evidence about the agency’s own cases. Are ownership rules stable? Does the message contain enough context? Are unusual but consequential requests present in the sample? A narrow router for routine messages may prove useful, but that would be a result to establish through the trial.
View image detailCheck a proposed tool call
Strands’ announcement describes a weather-tool example. A user asks about the weather without naming a city. An eager agent proposes a call with a city anyway. Before execution, the example asks whether the argument values are grounded in what the user said and whether the call is premature. In the illustrated path, the agent asks which city the user meant. Strands says the questions, threshold, and policy were chosen by hand for that demonstration.
A business workflow could ask a related question about a proposed account identifier. Did it come from information the workflow is authorized to use, or was it guessed? A reviewer would need the proposed call and the relevant evidence. Finding the same number somewhere in a conversation may be insufficient when several accounts appear. A later correction may supersede an earlier value. The rule must say which facts count.
Both error directions matter. Passing an invented value could cause a wrong lookup or update. Stopping a valid call could delay help and add review work. A contained trial would include corrected values, quoted values, and facts supplied through another authorized step. The weather example shows where a check can sit. It does not establish that its hand-picked policy works for another organization.
View image detailResolve an ambiguous payout problem
Imagine a customer says a payout went to the wrong destination, names two accounts, and asks the agent to “fix it today.” The next step may depend on account ownership, transaction state, permissions, and facts outside the message. The request contains several jobs: triage, verification, investigation, and perhaps a later change to a record.
A single “approve or deny the fix” choice hides those dependencies. Reviewers may be unable to establish an answer from the message alone. A wrong automatic action could be hard to reverse. A high reported confidence score cannot supply missing records or authority to act.
The agency can separate the work instead: identify an investigation owner, gather facts through authorized steps, clarify the account and transaction, and retain review for a consequential action. A bounded routing question might still merit a trial. Deciding the outcome of the case requires a different evidence base.
- Route a routine request: Information a reviewer needs: Message and written ownership rules; Consequential mistake to inspect: Delay or repeated handoff; Fallback to define: Ownership review for unclear cases
- Check a tool argument: Information a reviewer needs: Proposed call and authorized facts supporting its values; Consequential mistake to inspect: Unsupported value passes, or valid call is stopped; Fallback to define: Clarification or review before execution
- Resolve an ambiguous payout problem: Information a reviewer needs: Account, transaction, permission, and investigation facts; Consequential mistake to inspect: Wrong consequential action; Fallback to define: Gather facts and retain human review
These hypothetical cases differ in available evidence, error consequences, and fallback work. Those differences determine whether a local trial can answer a useful question.
View image detailRead the published results within their limits
The Strands repository reports 167 correct answers out of 231 public JevBench tasks at a 3,072-token window for v19. It also reports a median of 115 milliseconds per JevBench question on an RTX 3090 under WSL2. These are vendor-reported results under specified conditions. They do not measure the accuracy or total time of an agency’s routing workflow. The announcement’s latency graph caption refers to v18, so that graph should not be treated as a v19 measurement.
The captured sources disagree on some precise results. The repository lists a v19 Brier score of 0.342 and says v19 scored 168 of 231 at a 4,096-token window. The model card lists Brier scores of 0.348 at 4,096 tokens and 0.349 at 3,072 tokens, with 167 of 231 correct at either window. The captures do not reconcile those figures or establish identical evaluation settings for every Brier value. A precise comparison would require the underlying runs and settings.
The model card also identifies long, multi-step documents and transfer to unfamiliar yes/no questions or ordered rubrics as limitations. A short routing question may be worth testing, but its resemblance to a benchmark task does not prove local accuracy. A grounding check may look binary while requiring careful reading of a conversation and a proposed action.
View image detailThe relevant question is whether a particular decision, with its actual input and fallback, improves the existing path without introducing unacceptable mistakes. The captured sources establish no savings, safe threshold, or improved outcome for a reader’s workflow.
Run one contained trial
Choose one recurring decision, such as routing payout messages among established queues. Freeze its options, the information available at that point, and the action tied to each answer. Include a route for cases that cannot be assigned from the available facts. Keep later tool calls under their existing rules.
Assemble representative cases from work the agency is permitted to use. Include routine messages, confusing combinations, and rare cases whose mishandling matters. Ask responsible reviewers to apply a written rule. Preserve disagreements and their reasons; they may reveal a rule that needs repair before model evaluation.
Before examining model results, write down what would make the trial useful and what would stop it. The agency might identify an error it cannot allow in an automatic route, a maximum fallback load it can staff, and case types that always require a person. These are operating decisions for the agency. The captured Strands pages provide no validated threshold for them.
Run the proposed local path and the current path on the same permitted cases while the proposed path has no authority to take consequential action. Compare destinations case by case. Group mistakes by type and consequence. Examine correct and incorrect answers across confidence ranges, particularly near any threshold being considered. Count fallback cases and the time needed to resolve them. This is a proposed test; it was not performed for this article.
View image detailKeep the baseline visible. If people already route routine messages promptly, the local path must justify its setup and oversight. If the current path causes repeated handoffs, check whether the new route reduces them or moves work into ownership review. A public benchmark cannot answer either question.
Measure the whole path: time from available input to a usable action, compute demand, setup and maintenance, and human review effort. A fast model response may lead to enough clarification or review to erase a time advantage. Lower compute cost would be a poor trade if consequential mistakes increased. The captured sources contain no price or total operating-cost comparison for this workflow.
For an HTTP experiment, the v19 model card says the documented server binds to 127.0.0.1 and has no authentication. It describes local experimentation. A shared or public service would need a separate access and security design before it carried real workflow decisions.
At the end, compare results with the rules written before the trial. Identify which cases, if any, could take an automatic route, which need clarification, and which remain with a person. If the agency cannot make that distinction, the trial has found a workflow-definition problem worth fixing first.
Strands Decider 2B provides a way to test supplied-option decisions locally. Choose one decision whose answers, evidence, error costs, and fallback can be written down. Let reviewed local cases determine whether it belongs in an automatic path.
Checked for this article



