AI in Practice
GPT-6.1 Sol Cost Accounting: Caches, Retries and Review
A dated guide to GPT-6.1 Sol request accounting: price cache writes and reads correctly, apply the long-context rule, and attach retries and review to accepted work.

On this page
- Give the cost a finished unit
- Use the September 30 rate card as a ledger
- Price one request before scaling it
- Pay for the first cache write and count the reuse
- Preserve the prefix that can actually match
- Check the whole-request long-context rule
- Keep retries and tool calls on the same task
- Value review without pretending it is an API charge
- Compare configurations on the accepted result
- Turn the estimate into an accountable budget
To account for GPT-6.1 Sol agent work, reconcile fresh input, cache reads, cache writes and billed output for every request. Attach tool charges, retries and review time to the same task. Then calculate the provider bill and operating cost per accepted result.
This guide focuses on Sol's billing mechanics: a cold cache write versus a warm read, explicit prefix boundaries, full-request long-context pricing and a request ledger that preserves recovery costs. Our broader Rise comparison of API rates and accepted-task costs covers the cross-model candidate question. Here, the job is to explain where one Sol workflow's money and review time go.
OpenAI's GPT-6.1 Sol announcement compares its standard input and output prices with Astra's, describing them as one-fifth as high. That is a comparison of named token rates. It does not promise that a completed agency task costs one-fifth as much. This guide uses rates checked on September 30, 2026, and hypothetical calculations, not a performed Rise benchmark or a customer invoice.
Key Takeaways- Keep ordinary input, cache reads and cache writes in separate billing categories.- Price the full request under the long-context rule when input exceeds 272,000 tokens.- Keep failed attempts in the cost total and unresolved work beside the accepted-task count.- Measure review minutes before calling a token saving a business saving.
Give the cost a finished unit
Name the deliverable before you estimate the model bill. For a hypothetical agency creative brief, acceptance might require an approved audience, all requested sections, source-backed client facts and a recommendation that respects the agreed scope. A tidy document with an unsupported budget figure fails that rule. Its model charges still exist.
The creative brief is the task; initial generation and validation are steps within it. A retry repeats work to recover from a failed result. A successful first run can still contain several model calls. Give each task an ID and keep all its steps and retries together so you can explain its total.
Report the provider bill and the operating estimate separately. API cost per accepted task = total attributable model and tool charges / accepted tasks. The operating estimate adds explicitly valued review, repair and allocated maintenance before division. Showing both lets an owner see the cash charge and the staff capacity needed to finish the work.
Keep accepted, repaired, escalated and unresolved outcomes separate. If a handoff counts as completion, define its required contents. If a person finishes the brief manually, attach that labor before counting it as completed. The task ID must follow the work through recovery.
The ledger should identify the step to change. Extra context, repeated retrieval and missing source references call for different repairs.
Use the September 30 rate card as a ledger
The checked OpenAI model reference lists these GPT-6.1 Sol Standard API text rates in US dollars per one million tokens:
- Ordinary uncached input: Rate per million tokens: $2.00
- Cached input reads: Rate per million tokens: $0.10
- Cache writes: Rate per million tokens: $2.50
- Billed output: Rate per million tokens: $10.00
These are API rates. They do not calculate a ChatGPT or Codex subscription bill. For a subscription workflow, use its actual plan charges and applicable usage terms. An API model rate is a different purchasing unit, even when the selected model has the same name.
A token is a unit the model processes or generates. Budget from the usage record instead of treating a word count as the billable total. In particular, OpenAI's reasoning guide says reasoning tokens are billed as output even when they are not visible in the API response. The short answer a reviewer sees may be only part of the output charge.
For this guide's Standard, short-context calculations, the token charge is: ordinary input millions multiplied by $2, plus cached-read millions multiplied by $0.10, plus cache-write millions multiplied by $2.50, plus billed-output millions multiplied by $10. The input categories are mutually exclusive. OpenAI's caching guide explicitly says cache-write pricing is not an additive fee. Its usage calculation takes total input minus cache reads minus cache writes to obtain ordinary input. Apply each category's rate once.
Keep the rate date and processing mode beside that formula. The model reference also lists Fast at twice Standard, Batch and Flex at 50% below Standard, and a 10% regional processing premium where available. These are conditional terms. The examples below assume Standard processing without a regional premium, credits, negotiated terms or tax, and exclude tools until tools are expressly added. A different arrangement needs its own applicable-rate row.
Price one request before scaling it
Consider a hypothetical request that reads 40,000 fresh input tokens and generates 3,000 billable output tokens. Assume no caching, no paid tools and no retry. Under the checked Standard rates, input costs 0.04 million multiplied by $2, or $0.08. Output costs 0.003 million multiplied by $10, or $0.03. The model total is $0.11 for that request.
This stipulated usage is a request estimate. If the application sends the brand guide, prior discussion and retrieved documents, include all that input. A validation request has its own usage. Sum the requests attached to the brief's task ID.
Output volume changes the estimate. Suppose the same request generates 12,000 billed output tokens instead of 3,000 while input stays fixed. Output then costs $0.12 and the request totals $0.20. A verbose result is not inherently better or worse. The question is whether its extra material contributes to acceptance or merely gives the reviewer more to read.
Ask for the artifact the downstream process needs. A compact structured extraction and a developed client narrative have different output requirements. Making the extraction produce a long essay can add cost and obscure missing fields. Making the narrative too short can leave out the reasoning the client needs. A sensible output budget follows the deliverable rather than treating minimum length as a universal saving.
At scale, preserve the same units. One thousand identical $0.11 requests would produce $110 in stipulated model charges. They would not establish 1,000 accepted briefs. A forecast needs the expected number of steps, failed attempts and accepted results, along with the range of source lengths. Repeated source material creates a second billing question: how much is processed from scratch, written once or read again?
Pay for the first cache write and count the reuse
Imagine an application repeatedly using a 100,000-token approved brand library. Each request adds 10,000 fresh tokens and produces 4,000 billed output tokens. This hypothetical sequence assumes one cold write of the complete stable prefix, followed by nine full cache reads, all below the long-context threshold and with no tools or retries.
The cold request costs $0.25 to write the prefix, $0.02 for fresh input and $0.04 for output, totaling $0.31. Each later request costs $0.01 to read that prefix, plus the same $0.02 and $0.04, totaling $0.07. Across one cold request and nine warm requests, the model charge is $0.31 plus nine times $0.07, or $0.94. This counts ten requests, regardless of how many finished tasks they represent.
Without caching, each request would process 110,000 ordinary input tokens and the same output. That is $0.22 plus $0.04, or $0.26 per request. Ten would cost $2.60. The stipulated cache sequence saves $1.66 in model charges. It does not measure a production hit rate, accepted output quality or elapsed time.
The cache decision has a small upfront premium in this scenario. Writing the 100,000-token prefix costs $0.25 rather than $0.20 as ordinary input, a $0.05 difference. A full later read costs $0.01 rather than $0.20, saving $0.19. One reuse therefore covers the first-write premium under these assumptions. A prefix written once and never reused produces the opposite result: a higher model charge.
That gives the builder a more useful question than whether caching is switched on: how often does this exact prefix get reused after a write? Separate a reusable agency library from a one-off client's changing material. A library used throughout a working session has a different reuse opportunity from a document sent once and discarded. Do not assume the same savings for both.
Forecast cold starts as well as warm reads. Changed libraries can create many writes and few reads; several stable library versions can each incur a first-write cost. The reported categories show whether the intended reuse happened.
Preserve the prefix that can actually match
OpenAI's current prompt-caching guide requires matching rendered prefixes at eligible breakpoints. For developer-selected boundaries, mark the stable content with prompt_cache_breakpoint and set prompt_cache_options.mode to explicit when only those selected breakpoints should participate. A prompt_cache_key is optional for separate customer or user accounting; the current guide says these newer models handle routing automatically. In explicit mode, an unmarked request does not create cache reads or writes. The guide reports reads through cached_tokens and writes through cache_write_tokens.
For the creative-brief brand-library example, place the library and stable instructions before the boundary, then the current client material after it. A timestamp or case-specific detail inserted before that boundary changes the prefix you meant to reuse. Moving such content is useful only when the application still supplies the correct instructions and preserves the intended meaning.
Version the stable material deliberately. Suppose a team revises its approved messaging on Wednesday. New work should use the revised library, even if the change requires a fresh cache write. Continuing to send an outdated library for a cheaper cache hit would optimize the wrong result. Cache economy serves the content contract; it does not decide which instructions are correct.
Compare the first request's writes with later reads. If reuse is low, inspect the rendered prefix, including instructions and tool definitions. Correct the matching problem before treating the warm-request estimate as the normal bill.
A successful cache read tells you how input was billed. It does not establish that the creative brief passes review. It also does not establish that the request uses the base rates: a sufficiently large packet triggers a different pricing tier, even when much of it is cached.
Check the whole-request long-context rule
A long source packet can reprice the full request. OpenAI's GPT-6.1 Sol reference says prompts with more than 272,000 input tokens use twice the input and cache rates and 1.5 times the output rate for the full request. The example below applies that rule to the complete request, not just the portion above the threshold.
Assume 280,000 fresh input tokens and 8,000 billed output tokens, with no cache or other charges. Input costs 0.28 million multiplied by $4, or $1.12. Output costs 0.008 million multiplied by $15, or $0.12. The request totals $1.24.
For comparison, a separate 260,000-input-token request with the same output volume stays below the threshold. It costs $0.52 for input and $0.08 for output, or $0.60. The $0.64 difference is not merely the cost of 20,000 additional input tokens. The larger request changes the applicable rates across both input and output.
The two requests are not assumed to finish equivalent work. Removing 20,000 tokens may remove necessary evidence. First look for duplicated material or irrelevant history, then verify that the remaining packet still supports every required finding. Reducing input is useful only if the completion standard survives.
Cached input does not remove this threshold. A separate hypothetical 280,000-token request containing 240,000 cached-read tokens, 40,000 ordinary tokens and 8,000 output tokens would cost $0.048 for reads, $0.16 for ordinary input and $0.12 for output, totaling $0.328. No new writes are assumed. Most input is reused, but the long-context cache rate still applies.
Budget the source packet, not just its headline document. A workflow can accumulate earlier turns, retrieved text and tool results before a later request. Record input size per request so the team can see which step crossed the boundary. Possible repairs include selecting relevant sources, removing duplicate history or redesigning the sequence. Each repair needs an acceptance check and its own cost estimate because extra selection calls can add charges.
Keep retries and tool calls on the same task
Now return to the $0.11 request and build a hypothetical cohort of 50 briefs. Assume 40 pass on the first attempt, eight pass after one retry and two remain unresolved after that allowed retry. The ten briefs that did not pass initially each receive a second attempt. There are 60 attempts and 48 accepted briefs.
Assume each attempt uses the stipulated $0.11 token mix. Model charges total $6.60. Add 90 billable tool calls at an invented budgeting assumption of $0.015 each, or $1.35. This is not a quoted OpenAI tool price. A real ledger must use the charge for the actual tool, service, region and processing arrangement. The combined model-and-tool total here is $7.95.
Divide $7.95 by 48 accepted briefs to obtain about $0.166 in API and tool charges per accepted brief. The two unresolved briefs remain in the outcome report, and the charges for trying them remain in the numerator. Dividing by 60 attempts would answer a different question. Dividing by all 50 assigned briefs while calling the result accepted-task cost would misstate completion.
Tool use also explains why a workflow's calls may grow. An agent might retrieve a missing source, inspect a generated file and request another model step to interpret the result. Those may be necessary completion steps. Repeated calls that yield no new information need a different recovery rule. Log the purpose and result. A missing source may justify another retrieval; repeating a failed retrieval without changing the query or route points to a recovery problem.
Set a recovery budget before the run. For example, a proposed brief workflow could allow one correction when a validator finds a missing field, then send unresolved cases to a person. The exact policy depends on the task's consequences and deadline. Cost limits should identify when to change routes or request review, while preserving the actual completion state.
Price every retry from its own usage. A correction may resend history, add source material or generate a longer answer. The equal-$0.11 attempts simplify this example; a real recovery step can have a different token mix and cross a different pricing threshold.
Value review without pretending it is an API charge
Assume the hypothetical cohort consumes 120 minutes of human review and repair in total, including attention spent on the two unresolved briefs. At an assumed internal labor value of $45 per hour, that time costs $90. Add the $7.95 in model and tool charges and the modeled operating total is $97.95. Divided by 48 accepted briefs, that is about $2.04 per accepted brief, with unresolved work still reported separately.
The $45 figure is a scenario input. Choose an explicit basis, such as fully loaded employee cost or contractor charges. Saved review minutes may create capacity, reduce overtime or free attention; they do not automatically reduce cash payroll.
The separate components also reveal a false economy. Suppose a shorter source packet lowers model charges for the same cohort by $1.80, while tools and accepted count stay fixed. If it also adds an hour of review, the assumed labor cost increases by $45. The modeled total rises by $43.20. That outcome is hypothetical, but it explains why token savings alone cannot settle the decision.
Separate verification from repair in the review record. Checking a correct source citation may be part of the intended job. Reconstructing a missing citation is recovery work. A route that removes routine checking but creates more reconstruction may shift attention onto a harder task. Record what the reviewer did and why, then choose a repair: clearer required fields, better source references or a different model configuration.
Keep waiting time beside review minutes rather than silently charging both as labor. A request can take longer while a reviewer does other work. If the wait blocks a live customer or prevents delivery before a deadline, that delay has a distinct operating consequence. Record elapsed completion time and any actual blocked staff time separately before assigning a monetary value.
Compare configurations on the accepted result
Reasoning settings belong in the ledger. The checked model reference supports low, medium, high, xhigh and max effort, with medium the default. None and minimal are unsupported. Keep the selected setting with every attempt so a change in billed output or acceptance can be investigated against the configuration actually used.
Independent evaluation can help frame a test without replacing the invoice. Artificial Analysis's GPT-6.1 Sol release page defines its cost per Intelligence Index task as a weighted evaluation measure that includes input, cache hits, cache writes, reasoning and answer tokens. That is a useful example of counting more than the visible response. Its benchmark tasks and weighting do not price your agency's briefs, tool contracts or reviewer attention.
Compare configurations with the same acceptance rule and representative source packets. Record billed output, review minutes, final acceptance and elapsed completion time for each. A higher-effort route earns its additional consumption only if it improves a result the team needs. A lower-effort route earns its saving only if the accepted-task record supports it.
Use separate workload groups. A short brief with one clean source and a long brief with conflicting versions can have different cost drivers. Averaging them may hide a high-cost exception group. A team might keep one default and a documented exception route, but extra routing and maintenance must justify themselves. Adding a complicated classifier to save a few cents is another cost decision.
Keep the invoice explanation with the configuration decision. The team should be able to say which changes altered ordinary input, cache reuse, billed output or review, then check whether the accepted result justified the total.
Turn the estimate into an accountable budget
Start a proposed trial with one deliverable family and its existing completion standard. Select ordinary cases and known exceptions. Preserve their source packets so each configuration receives equivalent material. Assign a stable task ID, and attach every model attempt, tool call and review decision to that ID. This is a proposed measurement process, not a trial Rise has performed.
For each request, retain its model, effort, mode, rate date, full input count, ordinary input, cache reads, cache writes and billed output. Record the selected threshold tier and any applicable premium. Keep actual tool charges alongside the token calculation. For each task, record attempts, review minutes, completion time and final status. These fields let a builder explain a total instead of asking the owner to trust it.
Check a sample of the request ledger against the provider's reported usage and charges. Confirm that cache writes were charged once in their own category and that long-context requests received the full-request adjustment. Investigate unexplained differences before extrapolating the budget. Credits, invoice rounding and contractual terms should remain explicit adjustments, not silent reasons to change the underlying token counts.
Then forecast workload mix. Estimate how many tasks will use a cold prefix, reuse a warm one, exceed the context threshold or require a repair. Add maintenance as a separate line if it is part of the business decision. For example, a stipulated $300 monthly maintenance allocation spread across 960 accepted briefs adds $0.3125 per brief. If volume falls, that allocated amount rises even when the model rates stay fixed.
Set the spending limit and the decision before collecting results. A route needs to meet the completion standard, handle unresolved cases within the team's capacity and deliver at an acceptable operating cost. If it fails, the ledger should identify whether the next change belongs in caching, context selection, tools, output requirements or review. An estimate becomes useful when it tells you what to measure next.
Take one completed deliverable and reconcile its request records, tool charges and review time. Then repeat that check across ordinary work and exceptions. Choose a configuration when the accepted-result ledger supports its cost and the team can explain where the remaining work goes.
Checked for this article



