Skip to main content

AI in Practice

GPT-6.1 Sol: When to Use It for Everyday Agent Work

Use GPT-6.1 Sol where a bounded everyday task passes your checks. This guide develops document and coding trials, escalation rules and a reversible adoption decision.

OpenAI hands work to a person reviewing a document.
On this page
  1. The release changes the candidate list
  2. Confirm the surface and model contract
  3. Separate benchmark evidence from workflow fit
  4. Choose a task with a visible completion condition
  5. A document-to-delivery example
  6. A coding workflow with meaningful checks
  7. Keep escalation and permissions explicit
  8. Run a small, repeatable adoption evaluation
  9. Decide with evidence and preserve a return route

Test GPT-6.1 Sol on a recurring job with an inspectable finish: a source-backed document brief, a contained code change or a business workflow that produces a reviewable draft. Adopt it for that job only if it meets your current quality standard and keeps review manageable. The useful result is a handoff that lets a person make the remaining decision without rebuilding the agent's work.

OpenAI's September 29 announcement positions GPT-6.1 Sol closer to Astra on selected coding, computer-use and professional-work evaluations. As checked September 30, 2026, it lists ChatGPT Work and Codex access for Plus, Pro, Business, Enterprise and Edu users, with Chat still excluded. Developers can also use the API. Confirm availability in the account you'll use for the trial. OpenAI's launch announcement supplies the published access terms.

For an agency owner or operations leader, the opportunity is a narrower adoption decision than switching everything to the newest name. Choose a recurring task, describe what accepted work looks like, and find out whether this candidate can carry more of the routine steps without moving uncertainty onto your reviewer. The examples below are proposed workflows. Rise has not run a shared Sol-versus-Astra benchmark or measured customer savings for this article.

The release changes the candidate list

OpenAI reports improvements over GPT-6 Sol across software engineering, document questions, business automation and computer use. Its DeepSWE comparison reports a 6.4-percentage-point improvement over Sol's best score at lower reasoning effort and cost. Its AutomationBench comparison reports a 4.8-point gain at the same medium setting. OpenAI also says Astra retains the highest score in its difficult scientific-work evaluation. The announcement notes that research tools, prompts and settings can differ from production ChatGPT. Read the original evaluation context.

Rise reads these results as a reason to test GPT-6.1 Sol on substantial everyday work. Match the evaluation signal to the task: document interpretation can inform a briefing trial, while software-engineering results can inform a contained coding trial. Then check the candidate against your actual inputs and completion standard. Adoption evidence should describe the job you intend to change.

There is a practical difference between a candidate becoming worth testing and a live workflow becoming worth changing. Testing can be narrow and reversible. A default determines which route gets the next ordinary task, including tasks run by colleagues who never read the release announcement. That second decision needs a record a colleague can understand: the task class, the checks, the observed failures and the point where another route takes over.

This is why the release matters even if you have no reason to migrate today. It offers another candidate for work you currently send through a more expensive or more involved route. Keep the current process as the comparison. If you already have a reliable, inexpensive system for a fixed transformation, the new announcement may leave that decision exactly where it was.

Choose Actual size to read the graphic closely.

Confirm the surface and model contract

Confirm where the task will run before designing the trial. An API application, a Codex coding session and a ChatGPT Work task can surround the model with different tools and instructions. Access in one place doesn't prove that your existing connector or account exposes the same setup. OpenAI's Codex 0.159.1 release notes add GPT-6.1 Sol as the default in the bundled catalog. That is a catalog change, not evidence that every older installation or saved task changed its behavior.

For an API route, the current model reference identifies gpt-6.1-sol, supports text and image inputs with text output, and lists five reasoning settings: low, medium, high, xhigh and max. Medium is the default; none and minimal aren't supported. Tool calling uses the Responses API. These details were checked September 30, 2026 against the GPT-6.1 Sol model reference.

Write down the model, reasoning setting, application, tool access and source packet for each run. These details are the configuration being evaluated. If one person supplies files and another relies on a connector with different permissions, they have changed more than the model. Recording the setup makes a disappointing result easier to diagnose and a useful result easier to reproduce.

Check the intended output as well. A text response can describe a spreadsheet or propose a code change without proving that the requested file exists or that the change was applied. Decide whether the job ends in prose, a saved document, a proposed patch or an application state. Then make that destination part of the acceptance rule. The artifact should be inspectable where the next person will use it.

Choose Actual size to read the graphic closely.

Separate benchmark evidence from workflow fit

Independent evaluation helps challenge a vendor's selected comparisons, but it also has a defined scope. Artificial Analysis's GPT-6.1 Sol release page, inspected September 30, reports results by reasoning setting and describes an Intelligence Index built from ten evaluations. A composite tells you about performance across that collection. It doesn't measure your team's intake rules, reviewer preferences or delivery destination. Artificial Analysis's release analysis provides an independent evaluation view alongside OpenAI's launch claims.

Use the distinction to ask better questions. If your candidate task involves interpreting several approved documents, a document evaluation is a relevant signal. If it requires coordinating application steps, an automation evaluation is closer to the mechanism. Neither proves that an agent can reconcile two conflicting client instructions in your intake. That question needs an example with the conflict present and an acceptance rule that rewards surfacing it.

Reasoning settings belong in the comparison because the proposed setup is part of the product you are evaluating. Test the setting you intend to use for the task. If you change it, save the new result separately. Otherwise a successful high-effort run can quietly become the evidence for a medium-effort default that nobody actually checked.

Separate correctness, completion, review effort and elapsed time in the evaluation sheet. A brief can contain correct facts while omitting the required decision section. A patch can fix the requested behavior while demanding a long review of unrelated changes. A complete draft can arrive after the client deadline. Recording these dimensions separately shows which part of the job the candidate actually passed.

Give the reviewer the same rubric for each candidate. For a client brief, correct figures with missing source locations should receive the same disposition regardless of which model produced them. For a code change, a convincing explanation with an unrun relevant test should not receive extra credit because the release chart looks good. The rubric keeps the vendor's reputation out of the acceptance decision.

Choose Actual size to read the graphic closely.

Choose a task with a visible completion condition

Start with a task whose finish you can point to. "Help with operations" leaves the agent and reviewer guessing. "Prepare a draft exception report from the approved weekly records, with a source location for each finding and unresolved conflicts assigned to review" gives both of them something to inspect. A narrow job can still require substantial reasoning. Narrow describes its boundaries, not its difficulty.

Our practical test for what to automate considers repetition, input stability, judgment, failure consequences and maintenance. Apply those questions before comparing model names. A task that arrives in four different formats with no owner may need clearer intake first. A report nobody reads may need to disappear. A new model is a poor reason to preserve unnecessary work.

Choose a candidate with known inputs and a useful output. Weekly approved records, a fixed set of project files or a repeatable review packet are easier to compare than an open-ended request to discover everything important. Keep the recurring human decision visible: approving a recommendation, accepting a patch or deciding what to tell a client. The agent's task should prepare that decision with enough evidence to make the handoff easier.

  • Document briefing: Reviewable finish: Findings tied to approved source locations; Reason to hold or escalate: Conflicting instructions or missing evidence
  • Contained code change: Reviewable finish: Inspectable patch and relevant check results; Reason to hold or escalate: Unresolved design choice or failed required check
  • Business record update: Reviewable finish: Preview of exact records and proposed changes; Reason to hold or escalate: Ambiguous identity or a consequential external action

The same completion logic applies to all three rows, while the actual checks differ. A document source location cannot replace a software test. A passing test cannot authorize a customer commitment. Select a job where the person approving the result knows what evidence would change their decision.

Choose Actual size to read the graphic closely.

A document-to-delivery example

Imagine an agency preparing a weekly client delivery brief from an approved project-status sheet, meeting notes and a scope document. The requested draft must explain what is complete, what is delayed, what is blocked and what needs the client's decision. This is a hypothetical trial design, not a report from a performed GPT-6.1 Sol test.

Define a small output contract before running it. Each delivery item needs an owner, the recorded status, the relevant due date and a source location. Every recommendation must separate an observed record from an interpretation. Conflicting dates belong in an unresolved section. A missing owner remains missing. The agent is preparing the information for an accountable person, so guessing an owner to make the table look finished would fail the task.

Supply the same approved packet to your existing process and the candidate. Include the version date of each source so the reviewer can identify a superseded note. If the task includes a PDF, ask for source locations that a reader can follow in that actual document. Open the cited page or record and check the statement. If the link lands in a document but the claimed deadline appears nowhere on that page, the finding has not passed its source check.

Now examine a meaningful exception. Suppose the status sheet says a launch is Friday while the latest meeting note says the client has requested a hold. The acceptance rule should require the conflict to be surfaced with both locations. The proposed agent should not settle it by choosing whichever file looks more formal. That is a decision about which instruction governs, and your team needs an explicit rule or a human owner to resolve it.

Review the brief first for fields and source locations, then for the decision it prepares. A recorded delay should stay distinct from a recommendation to change scope. The client decision should be easy to find, and a blocker should remain clear even when the draft uses cautious language. This second pass checks whether the facts have been organized into a useful communication.

Save the defects in terms another reviewer can reuse. "Bad summary" is too broad. "The brief treated an unresolved hold as an approved Friday launch" identifies the issue, its evidence and its consequence. A specific defect can justify changing the source packet, the instruction or the route. It also helps distinguish a repeated reasoning problem from a single missing file.

The delivery step deserves its own check. If the job asks for a saved draft in a shared location, verify the file there and read the exported version. Review the content before sending it to the client. A capable document model can be worth adopting for preparation while the final communication remains a person's responsibility.

Choose Actual size to read the graphic closely.

A coding workflow with meaningful checks

The document trial checks a statement against its source. A coding trial checks a change against the behavior it must produce. Choose a contained change with a clear behavior. Imagine a reporting tool that exports dates inconsistently. The proposed task is to make the approved date format consistent in the export while preserving the surrounding record content. The agent can inspect the relevant files, propose a patch and run the checks appropriate to that change. This coding trial is hypothetical.

Write the before-and-after behavior in ordinary language. Give one representative input and the expected exported value. Identify any requirements the fix must preserve, such as the ordering of records or treatment of missing dates. If an unresolved choice affects the desired behavior, settle it in the brief or ask the agent to surface it before editing. Otherwise the trial tests its ability to invent your requirements.

Choose verification that would fail on the original date-export problem. Use the project's existing relevant tests, then inspect the exported values against the expected examples. A check that merely repeats the implementation can pass while the user-facing behavior stays wrong. The reviewer needs evidence about the requested result, including missing-date handling and record order where those are requirements.

Review the patch's scope alongside its results. A model may solve the date problem while changing unrelated formatting, introducing a dependency or cleaning up neighboring code. Those changes expand the review burden. For this trial, the target is a small, understandable fix. Record unnecessary edits as a separate defect so a technically correct outcome does not hide a costly handoff.

Also inspect how the agent reports blocked checks. If a required test cannot run because a dependency or service is unavailable, the final report should state the limitation. An explanation of why the patch looks correct isn't equivalent to a completed check. You can decide to review it manually or route it elsewhere, but the evaluation should preserve what was actually verified.

Compare candidate routes on the same inputs and working state. Reusing a project already modified by the first candidate can accidentally hand the next candidate part of the solution. Preserve the original task state for each run, retain the resulting patch and record the check output. This preserves the evidence needed to compare the routes.

The coding adoption decision can then be specific: use GPT-6.1 Sol for this class of contained maintenance work if its patches meet the behavior and scope checks with acceptable review. Keep architectural decisions or failures on the escalation route until they have their own evidence. A passing maintenance trial doesn't answer every software-engineering question.

Choose Actual size to read the graphic closely.

Keep escalation and permissions explicit

Write down where the candidate should stop. In the document example, the stop is an unresolved instruction or the delivery approval. In the code example, it may be an ambiguous requirement, a failed check or a change that expands scope. Stopping with a useful explanation can be a correct outcome. Forcing the agent to produce a finished-looking answer in every case rewards the wrong behavior.

Distinguish an escalation condition from a permission. Escalation asks who should handle the difficult part. Permission asks which actions the workflow is authorized to take. A stronger model may help diagnose a failed test, but that doesn't grant authority to deploy the fix. A better document draft doesn't grant authority to send it to a client. Keep these decisions in the workflow configuration and the operating brief.

When the stop condition is reached, hand over the evidence needed to continue. Include the task, work already completed, the unresolved issue and the decision required. In the document example, the handoff can attach the Friday launch record and the later hold request, then ask the owner which instruction governs. The reviewer should be able to resolve the issue without repeating the agent's search.

Make retries bounded too. If the candidate fails a check, one targeted repair may be a sensible next step. Decide the allowed attempt limit for the trial and record all attempts. An open-ended request to keep trying can consume the review window and obscure the first failure. The budget for persistence should reflect the task's deadline and consequence, not enthusiasm about agent autonomy.

Use a predictable return route. The existing process, an assigned reviewer or another evaluated model should be able to pick up the preserved work. Carry over the source packet and failure evidence, while identifying any assumptions that remain unresolved. Escalation is useful when it reduces the next person's search for context.

Choose Actual size to read the graphic closely.

Run a small, repeatable adoption evaluation

Choose examples that represent the work you actually intend to route. Include ordinary cases, recurring exceptions and one case where stopping is the accepted answer. For a brief, that might be a clean packet, a conflicting date and a missing required source. For the export fix, it might be a normal date, a missing date and a case requiring a product decision. Choose enough examples to examine your recurring failures; this proposed sample does not establish a general success rate.

Before viewing outputs, write the acceptance criteria and rejection reasons. Assign a reviewer who understands the job. Preserve each input packet, configuration, output and check result. Record reviewer corrections separately from the original result. Once a person has repaired an answer, it is easy to remember the corrected version as what the model produced.

Track accepted work, failures, stopping behavior, reviewer effort and elapsed time. Keep total spending available for the economic decision, but avoid letting a small invoice excuse a misleading brief or an overbroad patch. Calculate the economic case from its own usage ledger after establishing that the route delivers work the team can responsibly use.

Decide how to act on the possible outcomes. If ordinary cases pass and conflicts are surfaced reliably in your sample, you have evidence for a bounded next stage. If the outputs require frequent reconstruction, keep the current process while diagnosing the defects. If problems come from missing inputs, repair intake before attributing them to the model. If the candidate mishandles a consequential exception, retain that exception on the existing route.

Expand gradually enough that the evidence remains readable. Changing the model, source format, tool access and acceptance standard together makes the cause of a result difficult to identify. First establish the candidate under your intended setup. Then evaluate a meaningful change as a separate revision. The record should let a colleague reproduce the decision without attending your whole experiment.

Choose Actual size to read the graphic closely.

Decide with evidence and preserve a return route

Adopt GPT-6.1 Sol where a defined task passes your checks and the handoff is useful. Keep the existing route where the evidence is insufficient or where review absorbs the benefit. That can produce a mixed workflow: the candidate prepares a source-backed brief, a reviewer resolves client decisions, and an established coding route handles changes that need broader judgment.

Assign an owner to the new default and preserve the known-good configuration. Monitor the defect categories that mattered in the trial, especially hidden uncertainty, incomplete delivery and unnecessary scope. Revisit the route when sources, tools or requirements change. A model decision needs maintenance because the job around it can change even when the model name stays the same.

Pick one recurring task and write its trial brief: starting inputs, required result, checks and decision owner. Compare GPT-6.1 Sol with the process you already trust. If it passes, change that task's route and preserve the evidence. If it doesn't, keep the current route and use the defects to decide what to fix next. Either outcome should leave the team with a clearer decision than the release headline supplied.

Checked for this article

Sources

  1. OpenAI: Introducing GPT-6.1 Sol
  2. OpenAI: GPT-6.1 Sol model reference
  3. OpenAI: Codex 0.159.1 release notes
  4. Artificial Analysis: GPT-6.1 Sol release analysis

Keep going

All articles