Automation and Agents
Claude Opus 5.5 Pricing: What Does an Accepted Coding Task Cost?
Anthropic lists lower Opus 5.5 rates, but an agency’s useful measure is spend per accepted coding result. Here is a reviewable pilot and a transparent cost calculation.

On this page
- Define the result before pricing the run
- Read all four published billing categories
- Calculate an attempt, then include the failures
- Put review and repair beside the model bill
- Design a pilot that can answer your question
- Use launch results to choose candidates, not to fill in your results
- Check access, settings and routing before interpreting a result
- Make the decision from accepted work
Anthropic says Claude Opus 5.5 costs 40% less than Claude Opus 5 on typical workloads at default settings. If your agency is considering a long coding job, there is a more useful number to find: how much will you spend to get a change your team can accept?
Start with model spend per accepted result. Add the model charges from every planned attempt, including failed runs and retries. Divide by the number of results that pass an acceptance checklist written before the runs. Keep reviewer time, repairs, interventions and elapsed time beside that figure. A low token bill has limited value if the patch fails review or takes hours to untangle.
Anthropic’s 40% cost comparison is against Opus 5, not Claude Fable 5.1. Anthropic separately says Opus 5.5 performs at Fable 5.1’s level on most work. Both are company assessments. The prices and launch examples below come from the supplied capture of Anthropic’s September 22, 2026 announcement, captured September 29. Rise has not run the proposed comparison, checked access in an account or reviewed a bill.
Define the result before pricing the run
Suppose an agency wants to upgrade a dependency in a repository it is authorized to change. This is a proposed example, not a project we observed. Installing the requested version would be one requirement. The application would also need to preserve a named client-facing behavior, pass relevant tests and leave a diff a reviewer can understand.
For that task, an acceptance checklist might require:
- The dependency and lockfile show the intended version.
- The named behavior still works, including its most consequential edge case.
- Relevant tests pass from the same clean starting state used for every attempt.
- The diff contains no unexplained configuration changes, unrelated edits or new security exposure.
- A human reviewer accepts the result under the team’s normal process.
The real checklist should reflect the real job. A repository with weak tests may need a manual behavior check. A data migration may need a rollback plan and a comparison between old and new output. If the team cannot describe an acceptable result, a cost-per-acceptance figure will have an unclear denominator.
A generated patch is not automatically an accepted change. Tests can miss behavior nobody encoded. A diff can satisfy the prompt while changing an unrelated setting. A reviewer may find that a solution works today but would be difficult to maintain. Counting every generated patch as a success improves the spreadsheet without improving the software.
Record why each attempt fails. A failed test, altered behavior and an uninspectable patch all fail the checklist, but they point to different next steps. One may need a repair, another a clearer requirement, and another a smaller task. These notes can improve the workflow even if the pilot does not identify a preferred model.
Decide how repairs count before running anything. If a person fixes a patch before acceptance, record that time. If the agent receives feedback and tries again, include the added usage and intervention. If one configuration gets three repair attempts while another is judged on its first output, their results describe different opportunities.
For a repeated pilot on one task, each run should start from its own clean copy of the same commit. An accepted result means that candidate passed review; it does not mean the agency merged several copies of the same change. Keeping that distinction clear prevents a pilot’s accepted-count column from being mistaken for shipped work.
Read all four published billing categories
Anthropic’s September 22, 2026 announcement lists these US-dollar prices per million tokens for Opus 5.5 and Opus 5. They are published rates, not prices verified for a particular account, plan or cloud contract on September 29. Source: Anthropic’s pricing table.
- Input tokens: Claude Opus 5.5: $4.00; Claude Opus 5: $5.00
- Output tokens: Claude Opus 5.5: $20.00; Claude Opus 5: $25.00
- Cache reads: Claude Opus 5.5: $0.20; Claude Opus 5: $0.50
- Cache writes: Claude Opus 5.5: $5.00; Claude Opus 5: $6.25
These are separate kinds of billed usage. An agent that repeatedly works with repository context may incur cache reads. Building or refreshing cached context may incur cache writes. An attempt that produces extensive output may have a different bill from one that returns a compact patch. The mix belongs in the usage record. The size of the final diff cannot tell you what the model consumed along the way.
Use the categories as the applicable usage or billing record presents them. Do not count a billed cache read or write again as ordinary input merely because the material appeared in the model’s context. If a product surface reports charges differently, document that difference before comparing totals. Check whether tools or other services introduce separate charges under the terms that apply to the run.
Anthropic also lists fast mode for Opus 5.5 in Claude Code and the Claude Platform at $8 per million input tokens and $40 per million output tokens. It says fast mode offers up to 2.5 times the speed. That is a separate published input and output price and a vendor speed claim, not evidence that a particular job will finish 2.5 times sooner. The announcement does not specify fast-mode cache rates in that pricing sentence. Label fast-mode attempts and verify every applicable billing category before calculating their total. Source: Anthropic’s announcement.
The announcement does not provide a Fable 5.1 rate card. A three-model pilot needs the applicable Fable rate and billing record for the surface the team will use. A blank cell is more honest than a guessed price.
If a team already has a predictable token mix and a repeatable task, the published rates can support an initial budget. They cannot establish which configuration will produce the most accepted results or how much attention those results will require.
Calculate an attempt, then include the failures
For one attempt, multiply each billed token quantity, in millions, by its applicable rate. Add the resulting charges and any other documented fees that apply:
Attempt model spend = input charge + output charge + cache-read charge + cache-write charge + other applicable documented charges.
Here is an arithmetic example, not a usage estimate or a performed task. Imagine 10 million billed input tokens, 2 million output tokens, 50 million cache-read tokens and 2 million cache-write tokens. At Anthropic’s listed standard Opus 5.5 rates, those quantities would cost $40 for input, $40 for output, $10 for cache reads and $10 for cache writes: $100 total.
Hold those invented quantities constant and apply the listed Opus 5 rates. The charges become $50, $50, $25 and $12.50: $137.50 total. The $37.50 difference is about 27.3% of the Opus 5 total.
That rate-only example does not contradict Anthropic’s estimate of 40% lower cost on typical workloads at default settings. Anthropic says its estimate also reflects fewer tokens used per task. The example deliberately holds token use constant to isolate the published prices. It says nothing about how many tokens either model would use on a real job, whether it would finish, or whether a reviewer would accept the result. Source: Anthropic’s cost explanation.
Now include failed attempts. Suppose, only to illustrate the calculation, four attempts by an unnamed candidate cost $320 in total and produce one accepted result. Its model spend per accepted result is $320. Four attempts by a second unnamed candidate cost $500 and produce three accepted results. Its model spend per accepted result is about $166.67. These are invented outcomes for unnamed candidates, not Claude test results. The second candidate has the larger total bill but the lower spend per accepted result in this example.
Show the accepted count with the quotient. Three acceptances from four attempts convey something different from one acceptance from one attempt, even if the quotients happen to match. Also show spending on failures. If nothing passes review, report total spend and zero accepted results. There is no cost-per-success quotient when the accepted count is zero.
Decide where one attempt ends. If an agent resumes after a failed test, include the added usage. If the plan permits a fresh retry, retain the earlier run and its charge. If a reviewer comments on a patch and the agent revises it, record that exchange. Counting only the final successful run hides the cost of reaching it.
Set a time limit and retry cap. A candidate allowed to work indefinitely has a different chance to finish from one stopped after an hour. If the available surfaces make identical limits impractical, show the difference. The comparison may still inform a workflow decision, but readers of the results should know what was held constant.
Put review and repair beside the model bill
A cost metric is a way to make an operational decision clearer. The Work Worth Doing framework offers a broader lens for deciding whether a task is useful enough to automate and what responsibility should remain with a person.
The model bill is a charge. Review time is a staffing demand. Elapsed time affects delivery. Interventions show how much help the agent required. Keep these measures together on the scorecard without quietly converting them into one number.
In the proposed dependency upgrade, a reviewer might spend ten minutes checking the version and test result, then another hour tracing an unrelated configuration change. That is a hypothetical review path, not a reported Opus 5.5 outcome. The second hour would matter to an agency even if the token bill were small. Record what the reviewer actually had to do rather than treating every accepted patch as equally easy to inspect.
A useful record includes total model spend, accepted and failed attempts, elapsed time, reviewer minutes, repair minutes and interventions. Note whether existing tests made validation straightforward or the reviewer had to reconstruct the agent’s work. If a patch is accepted only after human repair, preserve that distinction. The team may still choose the workflow, but it should know which part of the result people supplied.
A team can put a dollar value on labor if it states the assumption. Hypothetically, 45 minutes of review at an assumed $80 per hour is $60. That is budgeting arithmetic, not a wage claim or an observed model result. Keep the assumed labor cost separate from the model charge so another team can substitute its own rate. If later rework becomes visible during the evaluation period, include it in the record too.
A more expensive attempt might finish sooner or be easier to review. A deadline-driven team might reasonably prefer it. The reverse could also be true. The published speed claim and a polished sample output cannot settle either question for this agency. The reviewer and delivery records from its own task must do that work.
Review capacity also limits the size of a sensible first trial. If only one overloaded engineer can judge a risky migration, a larger unattended run could produce a queue of plausible patches rather than usable progress. A bounded change lets the team learn whether it can confidently accept the work before it asks an agent to do more of it.
Design a pilot that can answer your question
Start with one authorized task that has a stable starting state and a clear acceptance checklist. A dependency upgrade or migration of one well-tested component could work. “Improve the app” is difficult to evaluate because almost any change can be presented as progress. A client-critical deployment without a review path or rollback plan is a poor first measurement.
Freeze the repository commit and create a clean starting copy for every attempt. Run baseline tests before a model changes anything. Write one task brief and one review rubric. Give each candidate the same task information and tools where the chosen surfaces permit it. Set the time limit, retry cap and number of attempts in advance. That prevents a disappointing run from quietly receiving an extra try after the team has seen the result.
For Opus 5.5, Opus 5 and, if relevant, Fable 5.1, record the exact model identifier, product surface, effort setting and cache policy. Match conditions where possible. If they cannot be matched, state the difference instead of describing the comparison as controlled. A Claude Code session, direct Claude Platform use and a cloud deployment may expose different controls or billing details. The practical choice is among configurations the team can actually use.
One run per candidate may reveal a feasibility problem, such as missing access or an unusable output. It is too small a sample for a broad reliability claim. More attempts can reveal variation, at the cost of more model spend and reviewer time. Choose a count that fits the decision and budget, then retain failures, timeouts and abandoned runs. Do not replace them with cleaner attempts after seeing the outcome.
For each attempt, save the prompt and configuration, starting commit, changed files, tests run, elapsed time, billed usage, retries, interventions and reviewer decision. Keep secrets out of comparison records. Have a reviewer apply the same checklist to every candidate. Reviewing without a model label may reduce expectations if the workflow permits it. If it does not, note that limitation. Either way, the reviewer must inspect the change itself.
Use the same columns for each configuration. Leave a value blank when it cannot be obtained, and explain why:
- Opus 5.5: record surface, effort and mode: Planned / completed attempts: Record; Accepted results: Record; Total model spend: Record; Failed-run spend: Record; Review / repair minutes: Record; Elapsed time and interventions: Record
- Opus 5: record surface and effort: Planned / completed attempts: Record; Accepted results: Record; Total model spend: Record; Failed-run spend: Record; Review / repair minutes: Record; Elapsed time and interventions: Record
- Fable 5.1, if included: record surface and effort: Planned / completed attempts: Record; Accepted results: Record; Total model spend: Record; Failed-run spend: Record; Review / repair minutes: Record; Elapsed time and interventions: Record
Keep the underlying per-attempt records. A summary table cannot explain an unusual failure or charge. After the planned runs, divide each configuration’s total model spend by its accepted count when that count is greater than zero. Read the quotient alongside the accepted count, failed-run spend, elapsed time and human effort.
Before choosing a winner, ask whether the measured difference is useful for this decision. A small pilot may show that one setup is feasible and another is not. It may reveal an obvious review burden. A narrow difference between two accepted-result costs, based on only a few attempts, may be too fragile to justify moving a whole codebase workflow. Record the observation and decide whether a second bounded task would change the decision.
The pilot may also reveal that the task needs work before any model comparison can answer the question. If every configuration fails the same hidden requirement, improve the brief or tests. If usage records do not expose comparable categories, state the limit on the cost comparison. If reviewers cannot confidently accept any patch, a lower rate card is not evidence for a larger unattended run. Each finding can guide a narrower next experiment without declaring a model winner.
Use launch results to choose candidates, not to fill in your results
Independent evaluations can add context, but their task definitions matter as much as their headline number. Artificial Analysis’s dated Opus 5.5 update places the model’s max-effort cost per Intelligence Index task level with Opus 5, despite Opus 5.5 generating about 1.6 times as many output tokens in that suite. Its broader model page describes a weighted average across ten evaluations. Those are measured benchmark costs under that evaluation’s setup, not an agency’s cost per accepted code change. The distinction is useful: a third-party “per task” figure still answers the task the benchmark defined, not necessarily the task a buyer needs to price. Artificial Analysis describes its Opus 5.5 results and scope here.
A separate coding-specific comparison, published by Bito on September 29, reports six in-house agentic coding tasks and repeated runs. It lists average session costs of $7.23 for Opus 5 and $2.71 for Opus 5.5, with average scores of 40.3 and 42.7 out of 50. The comparison used Claude Code defaults: Opus 5 at high effort and Opus 5.5 at medium. Bito also says Opus 5.5 served as its grader. That disclosed design makes the results useful evidence about one company’s task set, while the effort mismatch, model-based grading and limited six-task sample prevent treating the figures as a neutral universal ranking. Bito’s method and results deserve inspection before applying them to a different repository or acceptance bar.
Anthropic reports an internal exercise translating HAProxy from C to Rust with Opus 5.5 and Fable 5.1. According to the company, both rewrites passed nearly all HAProxy regression tests. Opus 5.5 finished in 9.5 hours versus 12 hours for Fable 5.1 and cost 51% less. That makes long coding work a reasonable category to investigate. It does not supply an accepted-result price for an agency’s dependency upgrade. The full harness, applicable billing details and human review effort were not independently checked for this article. Source: Anthropic’s coding section.
The announcement also reports 66.4% for Opus 5.5 on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 Main. Anthropic says most Opus 5.5 results in its table use adaptive thinking at maximum effort, while its Terminal-Bench result uses xhigh effort. Some comparison figures use different effort settings or come from another organization. Anthropic itself cautions that benchmark margins have become a less reliable guide to real-world differences at this capability level. These figures can support including a model in a pilot. They cannot replace that pilot’s task definition, acceptance checklist, usage record or reviewer. Source: Anthropic’s benchmark table and notes.
Anthropic describes one early tester completing a 680,000-line migration in under a day and a separate tester auditing and fixing a 200,000-line codebase in under three hours. Those are attributed examples under conditions this article cannot turn into a repeatable procedure. They do not show that every long task should run overnight or that every resulting patch will be inexpensive to review. Source: Anthropic’s announcement.
The examples and benchmarks provide reasons to investigate Opus 5.5. An agency’s repository, task boundaries, acceptance criteria and available reviewer determine what its own pilot can establish. Those details belong in the plan before money is spent.
Check access, settings and routing before interpreting a result
In its September 22, 2026 announcement, Anthropic says Opus 5.5 is available on its platforms and through Amazon Web Services, Google Cloud and Microsoft Azure. It gives claude-opus-5-5 as the Claude Platform model ID. That is Anthropic’s launch statement, not confirmation of access, prices or controls in a particular account, plan, region or cloud deployment on September 29. Check the intended surface before budgeting a job or recording a walkthrough. Source: Anthropic’s availability section.
The model selected at the start may not handle every part of a longer workflow. Anthropic says most cybersecurity tasks are rerouted to Opus 4.8, while routine bug identification and fixes remain possible. Its benchmark note says safeguard interventions sent cybersecurity tasks to Opus 4.8 and biology and frontier model-development tasks to Opus 5. None of that establishes that the proposed dependency upgrade would be rerouted. It does mean a workflow entering those areas may involve a fallback model. Record routing when it is visible, and treat unclear routing as a limit on conclusions about behavior or cost. Do not assign a billing effect without an applicable record. Source: Anthropic’s safeguards section and benchmark note.
Anthropic reports strong results on its automated behavioral audit and prompt-injection tests. The announcement also says evaluations cannot reliably catch every failure before deployment. The supplied capture of the Opus 5.5 system card returned no readable text, so its detailed methods and limitations remain unchecked here. The safety findings remain attributed to Anthropic. For an unattended coding task, set permissions and scope for that job and require review before merge or deployment. Source: Anthropic’s announcement.
Keep settings attached to each recorded result. Standard mode, fast mode and a run with a safeguard intervention describe different conditions even when the original task brief is identical. Results can still inform the agency’s choice if those differences remain visible.
Make the decision from accepted work
If you want a wider starting point across current model prices, our comparison of GPT-6 Sol and Luna, Claude Opus 5.5, and Gemini 3.8 Flash covers the broader cost-selection question. This article narrows the decision to how an agency should calculate and review coding work that actually passes its acceptance bar.
Anthropic’s published rates and reported coding results give an agency a reason to evaluate Opus 5.5. They do not identify the cheapest successful workflow for that agency.
If the team runs the proposed pilot, compare model spend per accepted result alongside accepted count, failed-run spend, elapsed time, interventions and review or repair time. State the sample size and any conditions that could not be matched. A larger job makes sense only when the recorded outcomes and available review capacity support it.
If the team has no pilot results, its next step is smaller: choose one authorized, reviewable task; define acceptance; then verify access and applicable rates on the surface it intends to use. The rate card prices billed tokens. The team’s acceptance decision determines whether those tokens produced work worth keeping.
Checked for this article



