Skip to main content

Automation and Agents

Grok 4.7 Coding Agents: Cost per Accepted Task

A practical pilot for measuring Grok 4.7 on your coding work, with a fixed acceptance rubric, reviewer effort, and cost per accepted task.

A coding task passes through a developer review gate before acceptance, with the authentic SpaceXAI publisher mark.
On this page
  1. What Grok 4.7 changes, according to xAI
  2. What an independent benchmark adds
  3. Define “accepted” before the agent sees the task
  4. Build a ten-task pilot that represents the queue
  5. Record completion, not just model output
  6. Calculate the number that matches the decision
  7. Let the result choose the next small step
  8. Where benchmark charts still help
  9. A practical experiment card
  10. The question is not “Is Grok 4.7 good?”

This guide uses published vendor and independent benchmark data; Rise has not run a Grok 4.7 test. Every proposed pilot example below is hypothetical.

A coding model is worth adopting when it completes your team's real work to an acceptable standard at a cost and review burden you can live with. A benchmark can tell you where to look. A launch post can tell you what the vendor is emphasizing. Neither tells you whether your own bug fixes, tests, migrations, and documentation will pass review.

That is the practical question behind Grok 4.7's September 21 launch. xAI says it trained Grok 4.7 on harder tasks that take many hours, with more emphasis on self-verification and longer context. It reports stronger results than Grok 4.6 across several of its chosen benchmarks. The useful next step for an engineering team is not to crown a winner from those claims. It is to set up a small comparison that measures accepted work, including the human time needed to trust and finish it.

What Grok 4.7 changes, according to xAI

The launch message is more specific than “a smarter model.” xAI describes a larger base model than Grok 4.6, a longer reinforcement-learning run, and a harder training mix weighted toward problems requiring many hours. It also says Grok 4.7 is better at checking its work, handling longer context, and understanding its Grok Bot harness. The API documentation describes a 500,000-token context window, low/medium/high/xhigh reasoning effort, text and image input, text output, and access to function calling, web search, X search, and code execution tools. Those are vendor descriptions of design and capability, not evidence that an agent in your repository will finish a multi-hour task.

The public launch table makes a more mixed picture visible. xAI reports Grok 4.7 xHigh at 46.3% on CursorBench 4.0, compared with 40.4% for Grok 4.6 High. On Terminal-Bench 4.0, xAI lists 37.6% against 20.3%. But its own table also reports 57.9% for Fable 5.1 Max on Terminal-Bench, above Grok 4.7, while other rows move differently. These are vendor-published benchmark results with different model settings and evaluation designs. They are not a clean controlled test of your team's own harness, acceptance criteria, or current alternatives. The asterisk on the Grok 4.7 DeepSWE figure also matters: xAI labels that score as high effort. Match settings before interpreting a comparison.

The launch page describes the model as being at the frontier in price-performance on CursorBench. Its pricing summary begins at $2 per million input tokens and $6 per million output tokens. “Begins” matters because the developer documentation describes different rates for requests above 200,000 prompt tokens, as well as separate cached-input pricing. A long agent run can cross that threshold. A task can also generate a lot of reasoning and output. So an appealing list rate does not settle the economics either.

Evidence belongs in separate lanes. Vendor claim · Outside benchmark · Your pilot. View image detail

Choose Actual size to read the graphic closely.

What an independent benchmark adds

Artificial Analysis's Grok 4.7 page provides another view, using its Artificial Analysis Intelligence Index and task-cost method. On October 3, 2026, the page reports an index score of 46, 73.9 output tokens per second, and an estimated $3.74 per Intelligence Index task for its xHigh configuration. Its comparison summary rounds speed to 74 tokens per second. These are time-sensitive observations from a third-party benchmark page, not fixed product specifications. The page describes a ten-evaluation index, including agentic knowledge work, professional tasks, terminal work, and long-context reasoning. It also records 240 million output tokens for Grok 4.7 in the index evaluation, compared with an 81 million stated median for its comparison class. Artificial Analysis defines task cost using input, cache-hit, cache-write, reasoning, and answer token prices, divided across tasks and weighted by index weights.

That is more informative than the $2/$6 rate alone because it measures a workload and counts generated tokens. For a broader cross-model view of rates, caching and task-cost arithmetic, see Rise’s AI task cost comparison. This article keeps the question narrower: what does your coding work cost when acceptance and review are included? It is still one external evaluation suite, not a prediction of how much a specific team will spend. A benchmark's “task” is defined by that benchmark. Your issue tracker may contain small deterministic fixes, ambiguous tickets, unfamiliar libraries and tests that take time to run. It also contains internal conventions and approval steps that a public benchmark never captures.

Independent does not mean universal. The benchmark can help you form a hypothesis: Grok 4.7 merits a place in a test if your work includes longer coding or knowledge tasks and the supported harness, context and API options fit. It cannot tell you whether your own reviewers will accept the results. Keep the claim this narrow and the data stays useful instead of becoming a decorative leaderboard.

A benchmark is one bounded window. Task score · Harness · Review · Integration. View image detail

Choose Actual size to read the graphic closely.

Define “accepted” before the agent sees the task

Teams often count a model's first answer as success because the answer is easy to see. That is a poor measure for coding work. A patch can compile and still change the wrong behavior. A test can pass while omitting the regression that mattered. A tidy explanation can conceal a security issue or leave the reviewer to reconstruct every assumption.

Write a short acceptance rubric first. It should match the type of work being tested and stay stable for every model. For a small code change, ask whether it solves the issue and preserves expected behavior. Does it include a useful test, pass the agreed checks and follow local patterns? It should also avoid creating a security or data-handling risk. It can also record whether the agent used the available tools correctly and whether the reviewer had to rewrite the result. This is not a contest to invent the longest checklist. The rubric should capture what “ready for merge” means in the work you actually assign.

Separate hard gates from quality ratings. If a change leaks a secret, breaks the required test suite, or edits an unrelated production path, a high score for elegant prose should not rescue it. Conversely, a minor naming preference should not turn a sound patch into a failure. Make the lines clear before the pilot so reviewers are not quietly moving the goalpost for whichever output surprised them.

For a fuller framework for assigning decision rights and keeping review with the right person, see Claude Opus 5.5 for unattended coding. Have two reviewers independently score a subset of results. They should not know which model produced the patch if the outputs can be presented without model labels. They will not always agree, and that disagreement is useful evidence. If reviewers cannot apply the rubric consistently, your first problem is probably not model selection. It is the definition of acceptable work.

Set the acceptance test first. Behavior · Tests · Safety · Reviewer approval. View image detail

Choose Actual size to read the graphic closely.

Build a ten-task pilot that represents the queue

A useful test set should look like the next month of work, not a collection of puzzles chosen because a model has a famous score on them. Ten tasks are enough for an initial directional signal if they come from several kinds of work. Ten near-identical bug tickets are not.

One workable mix is four routine changes, three tasks with moderate ambiguity or multiple files, and three longer tasks where context, tool use or verification actually matter. The exact proportions should come from your backlog. A small product team might sample different kinds of work: a bug fix, form validation change, focused refactor or missing test. It could also include a documentation update, data transformation, API integration, compatibility-sensitive migration, incident follow-up and a multi-step feature slice. An agency may choose client tasks that recur across accounts, but it must remove confidential data or use an approved environment before sending information to any model.

Choose tasks that can be evaluated against an existing specification, tests, acceptance notes, or a reviewer with domain knowledge. Do not use a task whose “correct” result is a personal preference unless you first record how that preference will be scored. Exclude work that could change live data or production systems during the experiment. The agent can propose a patch in a sandbox; people retain the authority to merge, deploy, or alter customer-facing records.

Capture the task's starting state. Keep the same repository commit, dependency versions, task text, supplied context, tool permissions, time budget, retry policy and test commands for each model. The harness is part of the result. If Grok gets repository search and shell tools while a comparison model receives only a pasted file, you have compared two packages, not models. If one run receives more retries or a much larger budget, note it and do not call the scores equivalent.

For tasks with multiple valid solutions, write down what counts as equivalent before reading the outputs. For a change with one expected result, preserve the acceptance tests but do not leak their answers into one model's prompt only. If you cannot reset the repository or test environment cleanly, run a single controlled sequence from fresh worktrees. Keep the human operator's hints in the record because a nudge can materially change the outcome.

Ten representative tasks, one test set. Routine · Ambiguous · Multi-step. View image detail

Choose Actual size to read the graphic closely.

Record completion, not just model output

For each attempt, record whether it completed, whether the patch passed the same automated checks, whether the reviewer accepted it, how much repair was needed, and whether the task was actually mergeable after review. If a model stops halfway, that is not a completed task even if its partial answer contains something useful. If the human reviewer finishes the hard part, count the human contribution rather than crediting the agent for the entire outcome.

Review burden is easy to underestimate. A model can produce a plausible patch quickly and still consume an hour of careful correction. Record active reviewer minutes, number of correction rounds, lines or files materially rewritten, and the reasons for rejection. Do not turn those fields into faux precision. Their purpose is to reveal whether “fast output” moved work from writing code to checking generated code.

Track failures by category. Examples include wrong interpretation, missed edge case, incorrect API use, insecure handling, tests that do not cover the behavior, unnecessary scope, tool-loop errors, stale context, or an output that never converges. A useful experiment finds where a model is a good fit and where it needs a smaller task, stricter boundary, or human lead. One overall score hides those operational differences.

The first run is not necessarily the final verdict. If one model fails because a task requires context it never received, record the failure. Then decide whether the context packet reflects a realistic workflow or a flawed test design. Do not retroactively add a helpful hint to one model's run and leave the other unchanged. If you repeat a task after repairing the harness, repeat the relevant models under the same repaired conditions and retain both results.

Record what the run actually did. Acceptance · Tests · Review minutes · Repair rounds. View image detail

Choose Actual size to read the graphic closely.

Calculate the number that matches the decision

The list-rate calculation is simple. Suppose a hypothetical short-context run uses 100,000 uncached input tokens and 100,000 output tokens. At xAI’s listed short-context rates, input costs $0.20 and output costs $0.60. The token total is $0.80 before account, tool, routing or operational charges. That is only a hypothetical calculation. It does not represent a tested task, a typical token volume, or the bill for every request. If a request reaches 200,000 prompt tokens, the long-context rates apply to all tokens in that request according to the pricing page. Cached tokens have a separate rate, and the amount that actually qualifies as cached input depends on the request and cache behavior.

Cost per accepted task is the combined cost of model tokens, directly metered tools or platform fees, and any reviewer labor your team chooses to price, divided by outcomes that meet your acceptance rubric. Keep token charges, review time and the combined result visible so the decision can be checked under different labor assumptions.

For example, a model could be cheaper in token charges but require more reviewer corrections. Another could cost more per run and still produce more accepted work per hour. Without your own task outcomes and a stated labor value, you cannot honestly decide which one is more cost-effective for your organization. If you do assign a labor cost, state the rate and what time is included. Keep setup, ongoing maintenance and evaluation overhead visible rather than hiding them in the token number.

Do not treat a vendor's benchmark cost-per-task and your pilot's cost-per-accepted-ticket as the same metric. Artificial Analysis publishes a useful defined benchmark estimate. Your calculation includes your task, your harness, and your acceptance rule. Label the two measures so a reader can understand why their values differ.

Count cost per accepted task: token cost plus tool fees plus reviewer labor priced by the team, divided by accepted tasks. View image detail

Choose Actual size to read the graphic closely.

Let the result choose the next small step

Do not start by asking whether Grok 4.7 should replace every other model. Ask whether it should get a narrow role in the routing policy. If it performs well on a class of straightforward, reversible tickets, that class can be the next supervised trial. If it handles long tasks but creates expensive review work, the team might reserve it for jobs where context continuity matters and tighten its stop condition. If it underperforms on a class, keep that work with the current process.

Decide the routing threshold before reading the results. For example, the team might require no critical safety failures, a minimum acceptance rate, and a review-time ceiling compared with its existing baseline. The threshold should reflect the cost of being wrong. A model helping draft internal documentation has a different risk profile from an agent editing payment logic or processing customer records.

Treat any early routing as reversible. Start with sandboxed work, one repository, a small task class, explicit tool permissions and human approval before merge. Set a run budget and a maximum number of tool steps. Keep a clear stop condition for unexpected file access, repeated failed checks, uncertain instructions or attempts to cross the task boundary. Longer-running capacity makes boundaries more important because the system has more opportunities to drift away from its assigned goal.

If the model's proposed change is good but the environment is a poor fit, you have learned something about your operating system, not merely the model. Maybe the task needs a clearer definition of done, a faster test command, a context file with current conventions or a safer way to reset state. Decide whether to fix that process for every model or make the model-specific routing explicit. Don't quietly tune the harness for one vendor and call the next score an even contest.

Route only the task class that passed. Keep other work on the current human-led path. View image detail

Choose Actual size to read the graphic closely.

Where benchmark charts still help

Benchmarks are useful for reducing the search space. A model that consistently struggles on the kind of work you need may not deserve a paid pilot. A model with results on long-horizon coding or agentic knowledge tasks can justify a test if those tasks matter to you. Benchmarks are also useful for seeing what a vendor chose to optimize and for tracking whether a release has changed relative to its predecessor.

They are less helpful when readers turn one score into a universal ranking. A coding benchmark captures the exact tasks, environment, scoring and model configuration that its authors selected. Your internal codebase, tool permissions, review norms and tolerance for risk differ. A leaderboard does not include all of the time a staff engineer spends reconciling unfamiliar output with internal constraints. It also may not capture a weak prompt, cache miss, orchestration issue or model outage in the way your production workflow will experience them.

Keep the chart next to its method. For xAI's launch comparison, label the source as vendor-reported and preserve effort settings. For Artificial Analysis, describe it as its Intelligence Index and note its task-cost approach. Don't combine figures from different benchmarks into one winner row. A percentage from one test is not on the same scale as an Elo score or dollars per benchmark task.

A procurement note can be decisive without claiming more than the evidence allows. For example: “Grok 4.7 has enough published evidence to earn a controlled trial for our longer coding tasks; we have not yet measured its success rate on our tickets.” That gives the team a next step. It does not confuse one external benchmark with a result it has not produced.

Read benchmark scores with their limits. Published score ≠ local acceptance rate. View image detail

Choose Actual size to read the graphic closely.

A practical experiment card

Before the pilot, record the repository and commit, eligible task classes, exact prompt and context, model and reasoning effort, and tool permissions. Add the time and token budgets, retry limit, required checks, reviewer rubric, accepted definition and stop conditions. This is enough detail for another teammate to understand what was compared without turning the exercise into a research program that takes longer than the tickets.

At the end of each run, save the complete result or a secure reference to it, tool-call log, test output, cost export, reviewer score and corrections. Follow your organization's data retention policy. Do not paste proprietary code into a public comparison spreadsheet or share unredacted prompts outside the approved access boundary. Treat model output as untrusted code. Review dependencies, scripts, tests and generated documentation before merge.

Summarize results by task class. A single mean can hide a model that is excellent on small fixes and unreliable on migrations. Show the numerator and denominator behind acceptance rate. “Seven accepted of ten tasks” is more useful than “70%” on its own, especially in a small pilot where one result moves the percentage by ten points. Describe how reviewers were assigned and whether they knew the model identity.

If two models are close, do not manufacture a winner. Consider the confidence you actually have, the differences in review effort, the platform features you need, and the cost of keeping two routing paths. If the evidence is too thin, the decision can be to run another bounded batch. A staged decision is still a decision. It is far more responsible than pretending ten attempts settled every codebase and task type.

Keep the first deployment reversible. Sandbox · Bounded tools · Human merge approval. View image detail

Choose Actual size to read the graphic closely.

The question is not “Is Grok 4.7 good?”

It is whether a particular version, reasoning setting and tool setup can handle a defined slice of your team's work with a level of correctness and review effort you are willing to own. xAI has published a clear bet on harder, longer coding and knowledge tasks. Its chart gives teams a reason to investigate. The independent Artificial Analysis evaluation adds another scoped signal, including token use and benchmark cost. Neither is a substitute for a carefully controlled sample of your own tasks.

If you are still deciding whether the task deserves automation at all, start with Rise’s Work Worth Doing test. Then choose ten examples from recurring, reversible coding work and write the acceptance rubric before the model sees them. Run Grok 4.7 and your current option through the same environment. Track accepted tasks, repair effort, reviewer time and total cost. Keep deployment and high-impact actions behind human approval. Then let the evidence decide whether the model earns a narrow role.

That is less cinematic than announcing a winner. It is also how an agent becomes a dependable tool: start with an honest boundary, measure what the team actually needs, and expand only when the next slice of work has earned it.

Checked for this article

Sources

  1. Introducing Grok 4.7, xAI, September 21, 2026
  2. Grok 4.7 API guide, xAI, updated September 28, 2026
  3. xAI API pricing
  4. Grok 4.7 model evaluation, Artificial Analysis

Keep going

All articles