Skip to main content

Automation and Agents

Claude Sonnet 5.5 Coding Agent Costs: Is Switching Worth It?

Anthropic claims up to 30% lower task costs, while one independent benchmark found higher costs at max effort. Compare the evidence and measure cost per accepted coding task.

Claude identity above a provider bill, a reviewed coding patch, and the handoff between them.
On this page
  1. What the launch establishes
  2. Keep the token rate separate from the task bill
  3. Work through the arithmetic before judging the claim
  4. Faster generation may or may not shorten the job
  5. Use the benchmark to choose a trial, not a verdict
  6. Choose work that can produce a useful answer
  7. Define an acceptable patch before either model starts
  8. Compare the same dimensions for both runs
  9. Keep configuration fixed enough to answer one question
  10. Confirm access before interpreting a missing result
  11. Record a safeguard fallback if it occurs
  12. Make a switch decision that fits the evidence

Short answer: Sonnet 5.5 is worth a bounded trial, but Anthropic's claim of up to 30% lower task cost does not tell a coding team what it will pay. The published input and output rates match Sonnet 5 at $2 and $10 per million tokens. The total also depends on output volume, cache use, retries, review, and whether the patch passes the same checks.

An independent September 28 evaluation complicates a simple cost-cut story. Artificial Analysis measured Sonnet 5.5 at $7.60 per task in its Intelligence Index at max effort, about 50% above Sonnet 5, alongside unusually high output-token use. The evaluation used a pre-release model with a structured-output issue Anthropic says it fixed for public release. Treat that as a benchmark signal, not your invoice. For a cross-provider example, see our dated GPT-6, Claude, and Gemini task-cost comparison.

What the launch establishes

Anthropic describes Sonnet 5.5 as a faster, lower-cost complement to Claude Opus 5.5. It identifies well-scoped everyday tasks and bug fixes as suitable work. Its announced API model ID is claude-sonnet-5-5. The Claude Code v2.1.284 release notes identify it as the default Sonnet model on the Anthropic API and list 1M context.

Those statements establish what Anthropic published. They do not establish that a particular team can select the model or use a 1M context window on its plan, provider, region, or configured product surface. A capability in a release note and access in your account are separate checks. The announcement is dated September 28, 2026, but the page does not establish a publication hour. The release page displays September 28 at 18:02 without a verified timezone. Neither detail supports a more precise launch timeline.

Anthropic also reports a large coding benchmark improvement. That is a reason to investigate the model, not a price for a completed task in your repository. The useful question is narrower: for work your team already delegates, can Sonnet 5.5 reach your acceptance bar with less total usage, elapsed time, or human correction?

Keep the token rate separate from the task bill

The rates in Anthropic’s September 28 announcement are per one million tokens. Keeping that unit visible prevents a claim about completed-task cost from being mistaken for a cut to the published price per token.

  • Input tokens: Sonnet 5.5 rate listed by Anthropic: $2 per million; What the announcement establishes about Sonnet 5: Anthropic says the rate is the same
  • Output tokens: Sonnet 5.5 rate listed by Anthropic: $10 per million; What the announcement establishes about Sonnet 5: Anthropic says the rate is the same
  • Cache reads: Sonnet 5.5 rate listed by Anthropic: $0.20 per million; What the announcement establishes about Sonnet 5: Anthropic says the rate is the same
  • Cache writes: Sonnet 5.5 rate listed by Anthropic: $2.50 per million; What the announcement establishes about Sonnet 5: The announcement does not give a Sonnet 5 cache-write comparison

The Claude Code release notes also list the Sonnet 5.5 model ID and the input, output, and cache-read rates. These are published terms, not a bill from a particular account. For a Sonnet 5 comparison, use the rates and usage categories that actually applied to its run. The reviewed sources do not establish its cache-write rate.

A rate is the charge for a unit of usage. A task bill reflects how much billable usage occurred while completing the work, along with any other applicable charges in the chosen setup. Fewer output tokens at the same output rate would reduce that portion of the bill. Another attempt could increase the total even if the first response was inexpensive. Human review is not an API token charge, but it is part of the team’s cost of delivering an acceptable change.

That leaves two questions worth recording separately. What did the provider charge for the attempts needed to reach an accepted result? How much review or repair did the result require? A team can later value staff time using its own assumptions. Turning an unmeasured review burden into a precise dollar saving would make the comparison look cleaner without making it more reliable.

Choose Actual size to read the graphic closely.

Work through the arithmetic before judging the claim

Consider hypothetical arithmetic, not a test of either model. One run uses 500,000 billable input tokens and 100,000 output tokens. At $2 per million input tokens, its input portion is $1. At $10 per million output tokens, its output portion is another $1. Its input-and-output subtotal is $2, before cache activity or other applicable charges.

A second hypothetical run at the same two rates uses 400,000 input tokens and 70,000 output tokens. Its input portion is $0.80 and its output portion is $0.70. The subtotal is $1.50, or 25% below the first subtotal. These invented counts show how a task could cost less without a lower published token rate. They say nothing about what either model actually uses, whether either patch is correct, or what a provider would charge after all categories are included.

For a Sonnet 5.5 estimate using the published rates, express each usage category in millions of tokens. Then calculate input × $2 + output × $10 + cache reads × $0.20 + cache writes × $2.50. Include other charges if they apply to your setup. Use the provider’s category definitions so a cached token is not counted twice. Keep the provider’s reported charge beside your estimate. If they differ, inspect the usage categories and configuration before announcing a saving.

The rates also show why token mix matters. At the listed rates, 100,000 output tokens account for $1, while 100,000 input tokens account for $0.20 before caching. A reduction in output can therefore have a different dollar effect from the same numerical reduction in input. That observation comes from the rate table; it is not evidence that Sonnet 5.5 will reduce either category on your tasks.

Now change the hypothetical outcome. If both first attempts pass the same tests and scope review, their $2 and $1.50 input-and-output subtotals can be compared for that narrow result, subject to omitted charges. If the $1.50 attempt fails and needs a second attempt with the same subtotal, its input-and-output charges reach $3 before cache activity, other charges, or review. The initially cheaper response has become the more expensive route in this invented example. The point is the finish line, not a prediction about either model.

Choose Actual size to read the graphic closely.

Cost per accepted task should include the attempts needed to meet the same standard. An unfinished attempt is still a cost, but it is not a successful task. If neither model reaches the bar within the agreed retry policy, the trial has no observed successful-task cost for that case. That finding may expose a weak task brief, an unsuitable task, or a model limitation. It cannot establish a cheaper winner.

Faster generation may or may not shorten the job

Anthropic’s 30%+ figure refers to output generation compared with Sonnet 5. It is not a measured 30% reduction in every coding task’s wall-clock time. A coding agent can spend time reading files, waiting for tools, running tests, revising an edit, or waiting for a person to review the patch. The share of time spent generating output depends on the task and setup.

Imagine two possible workloads. The first asks for a substantial patch and explanation in a repository whose tools respond quickly. Faster generation might shorten the agent run noticeably. The second has a slow test suite, repeated tool waits, and a brief final response. Generation could improve while the full job barely moves. These are workload illustrations, not observed Sonnet 5.5 results.

A useful timing record has at least two intervals. Measure from the start of the agent run to a result ready for review. Then measure review and correction until the change is accepted or rejected. If the tool exposes generation time, retain it as a third, narrower measure. Note unusual test waits and retries so a slow run is not automatically attributed to the model.

This matters when choosing a default. A developer may notice a faster stream and still wait just as long for the test suite. Another workflow may benefit because generation occupies much of its run. The published claim makes a trial reasonable. Task-level timing determines whether the speed change helps your work.

Choose Actual size to read the graphic closely.

Use the benchmark to choose a trial, not a verdict

Anthropic reports that Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5. It describes Terminal-Bench 4.0 as an agentic coding evaluation. These are vendor-reported figures. No independent reproduction is included in the reviewed source record for this article.

The difference is a reason to test a coding workload. It is not a forecast that Sonnet 5.5 will pass 70.6% of your bug fixes or deliver the same cost change in your repository. A benchmark’s tasks, environment, scoring, tools, and settings define what its score means. Your codebase has its own permissions, acceptance tests, and consequences when a patch changes too much. Anthropic’s announcement also shows that performance and cost comparisons vary by evaluation and effort level.

There is a fair counterpoint to a narrow local trial: one simple bug fix may understate the value of a model that handles harder work better. The reverse is also possible. A successful benchmark result may overstate its value for a team whose tasks are dominated by tool waits or demanding review. Neither possibility lets you skip measurement. It helps you choose representative tasks and keep the eventual conclusion within them.

The benchmark therefore helps choose the work to investigate. The next step is to select tasks the team actually delegates and define what an acceptable result looks like before either model begins.

Choose Actual size to read the graphic closely.

Choose work that can produce a useful answer

Start with a task your team might actually delegate again. A bounded bug fix is a good candidate when the failure can be reproduced, the permitted scope is clear, and a reviewer can judge the patch against written criteria. A task invented solely to make one model look good would not tell an agency much about its operating costs.

Avoid beginning with a request that has no agreed finish line, such as “improve this architecture.” That may be legitimate work, but its quality and review burden are harder to compare. A bounded version could ask for a specific change, identify affected interfaces, state what must remain compatible, and require a reviewer to confirm the result. The more open-ended the work, the more important it becomes to record judgment separately from test results.

Pick several tasks before interpreting the first run. They should represent the work that motivated the possible switch: perhaps a routine bug fix, a change with a slower test suite, and a task that requires reading more of the repository. Those are examples of selection criteria, not a claim that three tasks are statistically sufficient. If your team seldom delegates the third type, its result should not dominate your decision.

A single task can expose access problems, mismatched settings, or an acceptance check that needs work. It cannot establish a reliable percentage saving across a team. Repeated runs can help reveal whether a result was unusually lucky or unlucky, provided the team can afford them and keeps the task conditions clear. The practical goal is a defensible decision about a defined workload, not a universal model ranking.

Choose Actual size to read the graphic closely.

Define an acceptable patch before either model starts

Select a bounded bug with a reproducible failure. Save the repository state, such as the starting commit, and write an acceptance check before running either agent. The check might be a failing test or a repeatable incorrect output that should become correct. It should verify the requested behavior, not merely that a file changed.

Write down scope as well. Which subsystem may change? Which behavior must remain intact? Would a passing targeted test still be unacceptable if the patch edits unrelated code? A reviewer can then give a specific reason for rejection, such as “the test passes, but the patch changes an unrelated parser,” instead of relying on a vague impression of quality.

Use the same review procedure for both results. Reproduce the original failure, run the agreed check, inspect the changed files, and apply the same standard to out-of-scope edits. Set a retry limit before seeing either patch. Otherwise one model may receive extra chances because its first attempt looked promising, while the other is judged on its initial response.

Classify the outcome as accepted as submitted, accepted after a defined retry, accepted only after material human correction, or rejected. Record what the correction involved, not just how long it took. Ten minutes confirming a sound patch and ten minutes discovering a risky change have the same clock reading but different implications for future use.

If both first attempts pass, charges and timing can be compared directly for that task, with review differences attached. If only one passes, the failed attempt does not supply a successful-task price. If both fail, examine the task, acceptance criteria, prompt, tools, and outputs before blaming either model. A blank winner field is useful when the evidence has not earned a verdict.

Choose Actual size to read the graphic closely.

Compare the same dimensions for both runs

Use one row for each model run and the same fields for each row. Add an attempt number if retries occur. The following is a proposed record, not a completed trial.

  • Starting state and task: Record for Sonnet 5: Commit, prompt, acceptance check; Record for Sonnet 5.5: The same commit, prompt, and check; Why it matters: Confirms comparable work
  • Configuration: Record for Sonnet 5: Provider, surface, tools, effort, thinking, context option; Record for Sonnet 5.5: Corresponding settings; Why it matters: Reveals setup differences
  • Model actually used: Record for Sonnet 5: Requested and observed model, if exposed; Record for Sonnet 5.5: Requested and observed model, including any fallback notice; Why it matters: Prevents misattribution
  • Quality: Record for Sonnet 5: Test result, patch scope, review decision; Record for Sonnet 5.5: The same checks and decision; Why it matters: Defines an accepted result
  • Attempts: Record for Sonnet 5: Count and reason for each retry; Record for Sonnet 5.5: Count and reason for each retry; Why it matters: Captures work beyond the first response
  • Time: Record for Sonnet 5: Agent run, review, correction; Record for Sonnet 5.5: The same intervals; Why it matters: Separates generation from completion
  • Usage and charge: Record for Sonnet 5: Input, output, cache, applicable tools, provider charge; Record for Sonnet 5.5: The same categories; Why it matters: Explains the bill

Leave a field unknown when the product does not expose it. Unknown review time is not zero. An unavailable model identifier does not prove that the requested model handled the work. Label charges as reported charges or estimates. If a provider bills through a different arrangement, show that condition before comparing dollar figures.

Read the record in order. First ask whether both runs met the same acceptance bar. Then examine extra attempts and human correction. Finally compare elapsed time and charges for reaching that bar. This order prevents a low token subtotal from concealing an unfinished or risky patch.

Keep both the narrow API bill and the broader operating decision visible. One model could have a slightly higher provider charge but need less review. Another could produce a cheaper first attempt but require more retries. The right choice depends on which result the team can reproduce and which costs it can measure. Where the evidence is incomplete, describe the tradeoff instead of assigning a precise total saving.

Choose Actual size to read the graphic closely.

Keep configuration fixed enough to answer one question

For a model-focused pilot, start each run from a fresh copy of the same repository state. Supply the same task prompt, acceptance criteria, permitted tools, permissions, and time limit. Prevent the second run from seeing the first model’s patch or transcript. Save the inputs with the outputs so the team can later explain a difference in cost or quality.

Effort and thinking settings deserve attention. Anthropic says Claude Code and Claude apps default to Medium effort, while Claude Platform defaults to High. A comparison that silently inherits those different defaults changes both model and configuration. That may answer a legitimate deployment question if those are the setups the team would use, but call it a comparison of setups. If the goal is to isolate a model change, use comparable supported settings and record any difference that remains.

These are two valid questions with different answers. “Which model performs better under matched conditions?” helps explain a model change. “Which available setup should our team use tomorrow?” can include the settings, permissions, and billing arrangement the team would actually deploy. Decide which question you are asking before the runs. Otherwise a result can sound more general than the experiment allows.

Set a stopping rule as well. Decide how many attempts each model may make and what counts as material human repair. If a result fails those conditions, record the failure rather than prompting indefinitely until it passes. This keeps cost per accepted task from depending on unplanned generosity toward one run.

Choose Actual size to read the graphic closely.

Confirm access before interpreting a missing result

In its September 28 announcement, Anthropic says Sonnet 5.5 is available across its platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. The Claude Code release notes identify it as the default Sonnet model on the Anthropic API and list 1M context. Neither source verifies your plan, provider, region, account entitlement, or local configuration. Confirm that both models can be selected on the surface where the work will run. If a task needs a large context window, confirm that the relevant option is available there.

Save the requested model ID and, where exposed, the model that handled the request. Record the provider, product surface, effort, thinking setting, and context option. A display label may be sufficient for everyday use, but a configuration record is needed when explaining why two trial runs differed.

The announcement gives a migration condition: a workflow running Sonnet with thinking off must switch to the new between_tools setting before moving to Sonnet 5.5. Anthropic says that setting keeps upfront thinking off. Check whether this instruction applies to the workflow being tested. If migration changes its configuration, record the change instead of assigning every usage difference to the model.

An access failure is an operational finding, but it is not a performance loss for a model that never ran. Likewise, a claimed 1M-context demonstration requires an observed account and configuration. The published release note alone cannot supply that evidence.

Choose Actual size to read the graphic closely.

Record a safeguard fallback if it occurs

Anthropic says higher-risk cybersecurity requests can visibly fall back from Sonnet 5.5 to Sonnet 5. It also says routine software development is generally unaffected. Those statements do not predict how a particular prompt or account will be handled.

Preserve the actual model and any fallback notice in the run record. If a fallback occurs, describe that run as a fallback rather than treating its outcome as a clean Sonnet 5.5 result. Check the provider’s usage and billing information before assigning a cost to it; the announcement does not establish billing treatment for a specific fallback. If the model that handled a request cannot be confirmed, leave that comparison unresolved.

This matters when the pilot includes security-related coding work. The run may still teach the team something about the available workflow, but its quality and price should be attributed to the model that actually performed it.

Choose Actual size to read the graphic closely.

Make a switch decision that fits the evidence

The narrow decision is whether Sonnet 5.5 repeatedly completes a specified class of work to the team’s existing quality bar with less total time or charge under settings the team can actually use. Start with a bounded task, apply the matched procedure, and inspect accepted results. Repeat across representative jobs before changing a team default.

If both models pass and Sonnet 5.5 consistently uses fewer billable tokens or attempts, its unchanged listed input and output rates could still yield a lower bill. If it also reduces agent time without transferring correction work to a person, the operational case becomes stronger. If results differ by task, a limited routing rule may be more useful than one default for everything. These are conditional interpretations of evidence a pilot could produce, not outcomes observed by Rise.

A team may reasonably keep its current default if the newer model’s measured advantage is small, inconsistent, or outweighed by review work in the tested workload. That is not a contradiction of Anthropic’s claims. Its reported measurements describe its tests; a team’s decision concerns its own tasks, settings, and definition of acceptable work.

Postpone a broad verdict when access, effort settings, actual model identity, usage, charges, or patch quality remain unresolved. A trial can still be valuable in that state because it identifies the missing measurement. A published benchmark and a vendor cost claim are reasons to run a comparison, not substitutes for its record.

Save one task brief, starting state, acceptance check, settings, attempts, charges, and review decision for each model. If comparable results show a repeatable benefit on the work your team actually delegates, you have a reason to change the default for that work. Until then, Anthropic’s reported speed and task-cost gains are a reason to investigate, not a saving to put in your budget.

Checked for this article

Sources

  1. Anthropic: introducing Claude Sonnet 5.5Anthropic
  2. Anthropic GitHub: Claude Code v2.1.284 release notesAnthropic
  3. Artificial Analysis: independent Sonnet 5.5 evaluationArtificial Analysis

Keep going

All articles