Skip to main content

Automation and Agents

AI Agent Oversight: What to Check Before Giving an Agent Authority

Use Anthropic’s scoped oversight measurements to ask what an agent can change, when people review problems, and what evidence supports wider permission.

A human reviewer inspects a proposed document change while holding the control before it reaches the records archive, beside the authentic Claude logo.
On this page
  1. Define what the monitoring claim covers
  2. List the actions before choosing an approval point
  3. Put each review on the action’s clock
  4. Follow a failed check to a recorded decision
  5. Read block rates as activity counts
  6. Work one request through both paths
  7. Assemble an evidence packet for one permission
  8. Keep the report’s other measurements in context

Before an AI agent changes a client record, an operations leader needs to know what it can do, which actions are monitored, when a person can intervene, and what happens when a check fails. Give the agent authority only through the last step those answers support.

Consider a hypothetical request to update a client’s billing contact. An agent could read the account and prepare an edit. With broader permission, it could save the change, trigger another system to copy it, and tell the client the update is complete. Each step has a different consequence. Rise has not tested this workflow; it is an example for examining an authority decision.

Anthropic’s measurement report offers a useful vocabulary for that examination. It describes oversight measurements for Anthropic’s most-used internal research and engineering agent platform. Anthropic’s research listing dates the report September 17, 2026; the captured report text itself does not show a publication date. The reported figures describe Anthropic’s platform. An agency must gather evidence about its own workflow before granting permission.

The practical question is: What can this agent change, when can someone intervene, and what record will show the result?

The decision in brief: Trace one consequential action from the original request through the agent, monitor, destination system, and final human decision. Test known problems before widening permission. If the evidence covers preparation but not saving, let the agent prepare the work and hold the save for approval.
Approval before the save. Prepare the edit. Approve the proposed value. Only then allow the saved change. View image detail
The billing workflow and decision checkpoints are proposed examples, not tests performed by Rise.

Choose Actual size to read the graphic closely.

Define what the monitoring claim covers

Anthropic reports that approximately 30,000 research and engineering agents were active at any one time on its most-used internal platform as of August 2026. It says 100% of those agents’ actions passed through an online monitor before execution, usually within seconds, and 100% were ingested by an offline monitor afterward. The report limits these measurements to that platform. It does not say what share of all Anthropic agent activity the platform represents.

Coverage describes which actions pass through a monitor. It does not measure how often the monitor recognizes a harmful action. An online monitor can inspect and allow a mistaken action. An offline monitor can ingest a transcript without flagging it. Both actions still count toward the reported coverage.

In the hypothetical billing workflow, the agent platform might log a request to save an edit. The accounting application might keep a separate record of the value it actually saved. A notification service might send a message afterward. Monitoring the first system alone would not establish what happened in the other two.

Ask the team or provider to name the included agents, action types, tools, and destinations. Identify connected services outside that scope. Then follow one completed request: who initiated it, what the agent proposed, which check ran, what the accounting application accepted, and what the client was told. A trace that stops at the agent’s request cannot prove the final account state.

Follow the actual result. Agent and monitor records. Saved value in the accounting app. Notification and downstream copies. View image detail
The billing workflow and decision checkpoints are proposed examples, not tests performed by Rise.

Choose Actual size to read the graphic closely.

Record gaps separately. A monitor might see the proposed edit but miss the later message. It might record a successful tool call without retaining the changed value. Each gap calls for a different record or approval point. A single coverage percentage cannot resolve either one.

List the actions before choosing an approval point

“Update the billing contact” sounds like one task. In practice, the agent could read private account information, identify a new value, draft an edit, save it, and send a confirmation. Reading raises a data-access question. Drafting produces work someone can inspect. Saving changes the record. Sending a message may cause the client to rely on the change.

For each action, record the actor, permission, destination, and observable result. Include delegated agents and connected services. If another agent gathers the contact details, the record should connect that work to the original request. If the accounting application saves the edit, its record should show the value stored, not only the agent’s intention.

Name exclusions too. The agent might prepare a contact edit while lacking permission to change payment details. It might draft a client message but be unable to send it. These boundaries allow a useful limited role without bundling every billing action into one broad permission.

“Human in the loop” does not identify the boundary. A person might approve the initial request, inspect a proposed edit, check a saved result, or receive an alert after something goes wrong. Write the rule as an action: “A person must approve the proposed value before the accounting application saves it.” Then check whether a system permission enforces that rule. An instruction asking the agent to seek approval is weaker evidence than a control that prevents an unapproved save.

Permission is an action. Read and draft are distinct. Saving changes the account. Sending can make a client rely on it. View image detail
The billing workflow and decision checkpoints are proposed examples, not tests performed by Rise.

Choose Actual size to read the graphic closely.

Anthropic says its internal platform gives agents persistent identities and links shared messages to identities and transcripts. It presents these choices as ways to audit actions and follow communication among agents. That is Anthropic’s description of its platform. For a billing workflow, request the records that connect any delegated work to the final action.

A useful inventory pairs each proposed permission with the evidence needed to grant it. Preparing an edit may require a traceable source request and a reviewable proposed value. Saving it also requires evidence that approval works, the destination records the actual change, and someone can correct a bad result. The evidence burden grows with the consequence of the action.

Rise’s Work Worth Doing test helps decide whether a task merits automation in the first place. Once a team decides to delegate part of it, the action inventory defines how far that delegation goes.

Put each review on the action’s clock

Timing determines what a check can prevent. Anthropic distinguishes online monitoring before execution from offline monitoring after an action. Its report says the online automated check usually happens within seconds, and people review actions it blocks within one week. The offline monitor flags roughly 100,000 transcripts per week for further classifier review. Approximately 50 highest-priority flags per week are escalated to people. The report gives no precise time from an offline escalation to a human decision.

Those are different clocks. Seconds describes an automated decision before an action. One week is the stated human-review window for blocked online actions. The offline figures describe filtering and escalation volume. None establishes that a person reviews every action within seconds.

Apply those distinctions to the billing steps. A proposed edit can wait in a queue without changing a client record. A saved contact might be copied into another system. A confirmation might reach the client before an error is found. For each action, identify the first downstream effect and ask whether review can prevent it or reliably correct it afterward.

Request timestamps for the proposed action, automated decision, execution, later flag, assignment to a person, human decision, and correction. A fast classifier offers limited protection if its finding sits unassigned while the wrong value remains in use. Measure the path through a decision someone can act on.

Different reviews, different clocks. Online automated check: seconds. Blocked-action human review: within one week. Offline human decision: no precise time. View image detail
Anthropic’s reported figures describe its scoped internal platform; they do not measure this proposed billing workflow.

Choose Actual size to read the graphic closely.

“Reversible” also needs a practical meaning. The team might be able to edit an accounting value again but still need to correct copies and messages. A draft is easier to hold than a completed transaction. These consequences determine where an approval gate belongs; Anthropic’s internal review times cannot set that boundary for a client workflow.

Follow a failed check to a recorded decision

A block or flag starts a response. It does not tell the operator how the case ends. Anthropic says its online monitors can block or redirect actions before execution. It describes offline findings as inputs to investigation and improved mitigations. In a situation causing large imminent danger, Anthropic says it would shut down dangerous parts of an agent deployment. These are descriptions of Anthropic’s practices, not a response plan for an agency.

Suppose a pre-action check blocks the hypothetical billing agent from changing payment details. A useful case record confirms whether the account stayed unchanged, who received the case, and what they decided. It should also show whether the agent tried another route or a connected service sent a completion message despite the block. A block count cannot answer those questions.

An offline finding requires a different response. Suppose the agent saves a billing contact and a later review finds that the value conflicts with the authorized request. The reviewer needs the original request, saved record, and downstream action history. Someone needs authority to pause further edits, correct affected records, and decide when work may resume. This is a proposed response path; Rise has not observed such an incident.

Keep confirmed problems and false alarms in the record. False alarms consume review time; confirmed problems may justify narrower permissions. Close a case with the final account state and any unfinished correction. “Flag reviewed” only says someone examined a signal.

The reviewer also needs practical authority. An alert is of limited use if its recipient cannot inspect the source record, stop the next action, or authorize a correction. Name those powers before the workflow runs, then record who exercised them.

A flag starts a response. Block: confirm the action did not run. Later flag: trace what already changed. Give someone pause and repair powers. View image detail
The billing workflow and decision checkpoints are proposed examples, not tests performed by Rise.

Choose Actual size to read the graphic closely.

Read block rates as activity counts

Anthropic reports that its online monitor blocked 0.002% of more than a billion research and engineering agent decisions in August 2026. It also describes that rounded result as about one in 47,000. Its offline monitor flagged roughly one to two transcripts per thousand for further review. The measures use different units: online decisions and offline transcripts.

Neither rate counts all harmful attempts or establishes how often a monitor catches them. A low block rate could reflect rare prohibited attempts, the mix of work, the monitor’s threshold, or missed problems. The rate alone cannot distinguish among those explanations.

To assess detection in a particular workflow, define known problematic actions and expected responses before running tests. A billing team might try an unauthorized payment-detail change that should be stopped before saving. It might present a contact value that conflicts with the verified request, or try to send a completion message while the edit remains pending. These are proposed cases, not tests Rise has performed. Actual cases should match the system’s permissions and plausible failure paths.

For each case, record what was attempted, what the monitor could see, whether it blocked or flagged the action, whether a change took effect, when a person reviewed it, and how the case ended. Record misses and false alarms. Include delegated work if delegation is allowed and a downstream message if another service can send one.

Activity is not detection. 0.002% of online decisions blocked. 1–2 in 1,000 offline transcripts flagged. Different units; known-problem tests needed. View image detail
Anthropic’s reported figures describe its scoped internal platform; they do not measure this proposed billing workflow.

Choose Actual size to read the graphic closely.

One successful planted case shows how the system handled that case under those settings. It does not establish a general detection rate. A documented case can still reveal a missing log, a bypassed approval step, or a review that arrived after the client was affected. Use the finding to change a control or limit a permission.

Anthropic argues that developers should report how often monitors catch known agent misbehavior and have third parties test that independently. It says it plans to give external evaluators access to internal systems and data. The report mentions earlier METR red-teaming of its offline monitoring platform, but the captured material does not establish independent verification of all the figures discussed here or completion of the planned evaluator program. Ask a provider which tests were completed, who performed them, and what the reviewer could inspect.

Work one request through both paths

Take a hypothetical authorized request to change a billing contact. The record identifies the request’s source and the new value the client authorized. The agent reads the existing account and prepares an edit. If another agent gathered the detail, the trace shows that handoff. A reviewer can compare the proposal with the request before the save.

In the first path, the online check blocks the save. The trace records the attempted operation and monitor decision, then confirms that the application did not change the account. It also shows whether any connected service sent a message. A person resolves the case and records whether the edit was rejected, corrected, or authorized.

In the second path, the online check allows the save and offline review raises a concern later. The trace shows when the value changed and what happened before the flag. Did another system copy it? Was a confirmation sent? Can the team correct the account without leaving conflicting records? An isolated agent transcript cannot answer all of those questions.

Compare the paths at the same checkpoints: when the concern was found, what had changed, who could intervene, and how closure was recorded. If an error would affect the client before later review can act, hold that step for approval. If the effect is limited and can be corrected within a defined window, later review may support a specific, narrower permission. The actual workflow and its evidence decide.

Compare the same checkpoints. Blocked save: check the account state. Saved, then flagged: trace later effects. Both paths need intervention and closure. View image detail
The billing workflow and decision checkpoints are proposed examples, not tests performed by Rise.

Choose Actual size to read the graphic closely.

This walkthrough also exposes a communication gap. The agent might be barred from changing payment details but still able to send a message claiming those details were updated. Protecting the record would not protect the client from a false completion claim. Include both saved changes and messages in the action inventory and test cases.

Assemble an evidence packet for one permission

Tie the packet to an exact decision, such as whether the agent may save a billing-contact edit after preparing it.

First, request an action inventory and scope statement covering agents, delegated work, tools, destination systems, permissions, and exclusions. Add a trace of an ordinary completed request from source instruction to actual result. Handle sensitive client details appropriately while retaining enough structure to follow the action.

Second, request a blocked or flagged case through final disposition. Separate the automated pre-action decision, later automated review, and human decision. Include timestamps, the assigned role, pause and correction powers, final account state, and rule for resuming work. If only a constructed walkthrough is available, label it as a walkthrough. It explains the intended process but cannot establish how the process performed.

Third, request results from completed tests of known problems. Keep test plans separate. Ask how misses and false alarms were recorded and whether anyone outside the team that built the monitor inspected the outcomes. A plan helps decide what to test; it does not establish detection.

Finally, write down the permission decision and the evidence still missing. If the packet supports preparing edits but cannot account for saved values and downstream messages, grant drafting authority and hold saving and sending for approval. Revisit those boundaries when the missing traces, response times, and test results exist.

Evidence for one permission. Scope and a completed action trace. Failed check through final disposition. Completed tests, misses and false alarms. View image detail
The billing workflow and decision checkpoints are proposed examples, not tests performed by Rise.

Choose Actual size to read the graphic closely.

The voluntary NIST AI Risk Management Framework 1.0 Core addresses oversight roles, system context, risk measurement, monitoring, and response. It provides governance context for this packet. It does not validate Anthropic’s monitor or set an approval threshold for the billing example.

Keep the report’s other measurements in context

Anthropic also estimates that Claude “led” 26% of its measured AI research and development work as of August 2026, while no measured subset was rated fully autonomous. Its index weights task categories by estimated person-time and relies partly on judgments made by Claude. Anthropic describes limits to those judgments and to comparisons across labs. The estimate explains why Anthropic is asking for oversight measurements. It cannot justify a permission in a client workflow.

The report includes a separate compute-allocation snapshot for July 13–20. The captured report text does not explicitly identify that week’s year. Neither that snapshot nor the R&D automation index measures whether an agency can trace and correct a billing change.

For one proposed permission, map the actions, inspect a complete trace, establish the time through human disposition, and run documented cases for known problems. Grant authority through the last consequential step those results support. Keep the next step behind approval until the team can show what happens when it goes wrong.

For the report’s separate research estimate, read what Anthropic’s 26% AI-led research figure measures. Its task-category index answers a different question from whether an agent may change a client record.

Checked for this article

Sources

  1. Anthropic, “Measurements for understanding the pace of AI development inside frontier labs”
  2. Anthropic research listing
  3. NIST, AI Risk Management Framework 1.0 Core

Keep going

All articles