Skip to main content

Systems and Workflows

How to Set Safe Budgets for Exa Agent Ultra Research

Set an Exa Agent Ultra budget, duration and review gate with this practical preflight for safe handling of partial research results.

Exa Agent Ultra operations cover: a research request moves through a budget and time boundary into an evidence review, with the official Exa mark.
On this page
  1. Freeze the decision before writing the query
  2. Give the output a schema that makes uncertainty possible
  3. Choose the cost ceiling from the value of the work
  4. Pick a time limit around the decision deadline
  5. Make completion state part of the data model
  6. Validate evidence field by field
  7. A concrete preflight for a bounded research run
  8. Common setup mistakes
  9. A run should know how it can fail

Before an Exa Agent Ultra run informs an operating decision, define what the result must contain, what it may cost, and when it must stop. Decide in advance how a partial answer will be handled. Put those choices in the request and workflow instead of improvising after a large result arrives. Exa's current documentation describes Ultra as a metered, high-effort mode with a default $20 per-run cost cap and an optional duration limit from five minutes to three hours. Reaching a configured limit can leave a useful partial result. It does not turn that result into a complete list.

This article is an operational setup guide. It does not recommend Ultra for every task or claim that a well-formed schema guarantees correct research. The goal is narrower: make the run's boundaries observable, make each output auditable, and ensure a cutoff cannot quietly masquerade as successful completion. Product behavior and prices below were checked on September 30, 2026. No API run was executed for this guide.

Key takeaways- Define the task, acceptance rule, schema and unknown state before starting.- Set an explicit spend cap and duration based on the work’s value and deadline.- Route budget or time cutoffs to partial review, then validate claims against their cited pages.

Freeze the decision before writing the query

Choose Actual size to read the graphic closely.

Start with the decision the research will inform. “Find every relevant company” is a search request, but it is not yet a safe run definition. What counts as a company? Which geography, date range, industry boundary and evidence standard apply? Do products, divisions, acquired brands and parent companies count separately? If two researchers received the same brief, could they judge a row consistently?

Write the task's outcome before the agent's prompt. For example: “Prepare an initial diligence queue of companies that publicly document all four capabilities, with one direct source for each capability, as of the run date. Rows are leads for review, not approved vendors.” That statement separates two jobs. The agent assembles candidates and evidence. A person verifies whether a candidate may proceed. If the output will trigger a consequential legal or commercial step, define the responsible reviewer and the required approval gate before research begins.

Next, separate the finite from the open-ended. If there is a known universe, store its size or reference list. If the web population is open-ended, do not promise “all” as a result. Instead state what will count as a useful bounded run: a target number of candidates, a search boundary, a set of required sources, or a coverage check against a separately assembled sample. “Continue until exhaustion” describes the product's effort posture. It does not prove the public web was exhausted.

Write exclusions while the question is still clear. A run may need to omit known companies already in a CRM, disallowed jurisdictions, duplicate subsidiaries or unsupported product categories. Exclusions can be useful when expanding an existing list, but the rule should distinguish “already known” from “not eligible.” Save the seed list and the normalized exclusion rule. Otherwise, the team may later mistake a deliberately suppressed known item for a research failure.

Give the output a schema that makes uncertainty possible

Choose Actual size to read the graphic closely.

An output schema is a contract for what the agent should return, not an independent data-quality check. Use the schema to require fields that make the result reviewable. These may include a stable entity name, canonical domain or other identifier, the required classification, evidence URLs, excerpts or a short support rationale, an as-of date, and a confidence or review status. Keep each criterion as its own field so a reviewer can find which part failed.

Every field that can be unavailable should permit an honest unknown value. The Exa Agent quickstart advises making uncertain fields nullable and not required when the system may be unable to verify them. Its example distinguishes cannot_verify from an actual negative result. That distinction prevents a blocked site or an inaccessible document from being recorded as proof that a company does not meet the criterion. For a list-building task, you might use verified, not_verified, and needs_human_review states rather than forcing a binary pass when sources conflict.

Avoid schemas that demand unnecessary enrichment. A name, five contact details, funding history, technology stack and source for every field may make a task much more expensive and much harder to verify than the decision requires. Every required field adds a question the research must answer. Ask what the next person genuinely needs. Make optional context optional, and bound arrays or the candidate count when the output format allows it. That helps control size and reduce avoidable enrichment work.

Normalize dates and units in the description. “Recent,” “large,” and “nearby” are not stable fields without definitions. Prefer an explicit period, currency, geography, source date, or threshold. Require URLs to be canonical and evidence to be field-specific. If one landing page supports only one of four criteria, it should not be copied into all four evidence cells. A source list is not proof by proximity.

Choose the cost ceiling from the value of the work

Choose Actual size to read the graphic closely.

As checked September 30, 2026, Exa's Ultra documentation says metered Agent use is capped at $20 by default per run. budget.maxCostDollars accepts values from $1 to $100, but account or server limits may be lower. The cap is a maximum, not a fixed price. Agent compute, searches and other tools can contribute to use; a run that finishes early may cost less. Exa documents fixed per-request effort settings for more predictable pricing and says to choose one when predictable price matters. It describes Auto for variable scope and Ultra for exhaustive work where completeness matters more than speed or cost.

A safe cap is a business decision. If the task is exploratory, small and reversible, a modest ceiling may be enough to learn whether the prompt produces useful structure. If a missed entity could waste days of diligence or make a market map misleading, a larger cap may be reasonable, but only with an explicit research budget owner. Set the ceiling below the amount that would require a fresh approval. Keep the actual invoice separate from the cap in the run record.

Do not set the maximum simply because the API permits it. A $100 limit is not a recommendation or a prediction of quality. Before every run, estimate the maximum exposure across the number of runs, expected review time and any paid data tools. If you plan to test several effort settings on the same task, include the full comparison budget rather than evaluating each call in isolation. A good pilot should cost less than the decision it improves.

The important denominator is often not dollars per run. It is cost per accepted record or decision-ready result. A run that returns twice as many rows but requires three times as much human cleanup may be a poor exchange. Another run may cost more and still be worthwhile if its accepted additions change a consequential decision. Measure actual spend and correction labor under a written acceptance rule instead of treating a leaderboard's average task cost as your own economics.

Pick a time limit around the decision deadline

Choose Actual size to read the graphic closely.

The current Agent Ultra guide accepts maxDurationSeconds from 300 to 10,800 seconds, meaning five minutes to three hours. It says complex Ultra runs typically complete in about 30 minutes and difficult runs can take up to three hours. The docs list no default time limit. A run that has no explicit duration ceiling can therefore continue longer than an operational workflow expects, subject to service behavior and other limits. For an asynchronous process, set a limit that reflects the actual deadline and your tolerance for partial coverage.

The right duration is not automatically three hours. A legal diligence memo needed by a fixed meeting may have an hour available; an internal landscape review for next week may tolerate a longer run. A short cutoff can control latency while intentionally yielding a partial set. A longer duration can give the service more time to extend discovery. Neither setting establishes correctness. Document the reason for the chosen ceiling, and do not compare a run stopped after five minutes with one run to a longer limit as if effort were the only changed variable.

Plan the polling or streaming path too. Exa notes that SDK helper defaults may time out after one hour, while Ultra can run longer; callers should provide a longer client wait or consume events. An SDK timeout is not evidence that the server-side research failed or stopped. The workflow should persist the run ID as soon as it exists, track the remote status, and distinguish the local client stopping its wait from the Agent reaching its own time limit. If the execution environment can restart, store enough state to reconnect safely.

Be explicit about early cancellation. The documentation says stopping a run preserves collected results and bills usage to that point. That can help when a run is over budget in practical value or the request changes, but it still produces a stopped result, not a completed one. The downstream gate should route it to a partial-results lane. A human can decide whether that partial set helps; an automation should not silently remove its stop state.

Make completion state part of the data model

Choose Actual size to read the graphic closely.

A partial result is output that has not yet passed the request’s completeness test, whether a configured limit stopped the run or coverage remains unverified.

A row schema describes content. A run-state record describes whether the job is finished enough to act on. Keep these separate. For each run, save the request hash, effort, configured cost and duration caps, API run ID, start and finish timestamps, terminal status, stop reason, cost breakdown, returned output, schema validation result, evidence-check result, deduplication result and reviewer disposition. This provides an audit trail without embedding process metadata in every article or business record.

Exa's current Ultra guide lists several terminal reasons. schema_satisfied means the agent finished the task. budget_reached and time_limit_reached identify a limit. stopped means a caller stopped it early. error and cancelled identify unsuccessful terminal conditions. The docs say that approaching either configured limit causes the agent to stop starting new work and return what it found. Treat the stop reason as a material part of the result. Do not infer “complete” from an HTTP success, from status: completed, or from an output object alone.

A simple downstream state machine can keep those conditions visible: queued, running, candidate_result, partial_review, complete_review, rejected, and approved_for_next_step. A budget- or time-limited run goes to partial_review. An error or cancellation does not become an empty list. A completed schema-satisfied run still needs field checks. Only approved_for_next_step is eligible for whatever comes after research, and the human accountable for the decision should set or authorize that transition.

Retries need care. If a client loses its connection after creating a run, it should recover the existing ID before submitting another chargeable request. If a retry is required, label it as a new attempt and link it to the earlier run. Rise’s GPT-6.1 Sol cost-accounting guide gives a related example of keeping retries and review attached to the task; apply Exa’s own metered rates to an Ultra run. Deduplicate results across attempts with an explicit canonical identity rule, preserving both source evidence and the reason a duplicate was merged. A retry is not a continuation unless the API operation actually uses its documented continuation mechanism.

Validate evidence field by field

Choose Actual size to read the graphic closely.

An agent can return a well-formed JSON object with a wrong classification. The API's schema validation checks structure; it does not establish truth. For every consequential field, a reviewer or a separate validation step should open the cited source and check whether it supports the precise statement. Confirm the source's date, entity identity and relevant context. A vendor page can support what the vendor claims, but a claim about independent performance or comparative quality needs evidence suited to that claim.

Use a two-level review. First, check individual records: entity identity, criterion-by-criterion support, date, URL health and requested data shape. Second, check the set: duplicates, omitted categories, coverage across a known sample, and whether the list drifted into adjacent types of entities. One hundred verified rows do not demonstrate that no 101st exists. If the population is open, use language that describes the observed output and the reviewed boundary.

The WANDR paper illustrates why record and set review should stay separate. Its wide-and-deep benchmark checks each record and cited evidence, then aggregates performance across the requested hierarchy. A candidate may be partly correct while failing a required evidence field. A run may have clean sources for the rows it found but still miss much of the target volume. Conversely, increasing the returned count can increase incorrect or off-scope records. The research paper is a methodology reference, not evidence of how an untested Ultra run will perform for your team.

For a pilot, independently verify all high-risk rows and a stratified sample of the rest. Sample across categories, source types, regions, ambiguous matches and records that the system marked uncertain. If the team has an external reference list, use it to test recovery, but remember that it measures only what that list contains. If reviewers disagree, record the disagreement and revise criteria before the next run rather than counting whichever interpretation favors a provider.

A concrete preflight for a bounded research run

Choose Actual size to read the graphic closely.

Suppose a go-to-market team wants an initial list of software vendors that meet three public criteria. This example is illustrative, not a measured case. The team will use the list to choose which companies to research further. It does not plan to purchase or contact anyone automatically. That makes the result a triage artifact, not an approval record.

The preflight could state the source date, geographic scope, vendor definition, three criteria, disallowed substitutes, required evidence per criterion, identifier rule and unknown state. The schema would require a canonical domain, company name, three criterion statuses, one source URL and supporting excerpt for each status, last-checked date, and a short rationale. It would allow null plus cannot_verify for inaccessible or ambiguous claims. Known vendors already in the list would be supplied as exclusions, while a separate sample of known eligible vendors would be reserved for the evaluation rather than excluded.

The team could run a small test first, with a deliberately modest cost cap and a duration matched to its workday. It should inspect the returned schema, exact stop reason, charges, grounding and a sample of evidence. If the output is mostly strong but incomplete at the chosen cutoff, a second run with a changed duration may teach something. If the output confuses vendors with agencies, more time probably will not fix the definition; revise the task before spending again.

For a decision-relevant comparison, freeze a representative task and run both the current baseline and Ultra under comparable rules. Decide how to blind reviewers, who adjudicates disputes and what acceptance threshold counts as useful. Track raw candidates, unique candidates, fully supported candidates, partial candidates, unsupported claims, omissions from the reference sample and human minutes. Report the sample design alongside every score. Without that, a neat percentage can look more scientific than the evidence deserves.

After review, choose one of three actions: reject the output, retain it as an explicitly incomplete research lead, or approve a verified subset for the next human-controlled step. Do not let a strong-looking result bypass the state transition. The person who owns the business decision can approve a subset even when the whole run is partial, but the record should preserve that limitation and exactly which rows were approved.

Common setup mistakes

Choose Actual size to read the graphic closely.

Using “all” without defining the universe. The word sounds decisive but rarely supplies a measurable stopping condition. Define a known population or a bounded research objective. Record when coverage is not measurable.

Requiring every field to be present. A schema that cannot express unknowns nudges a system toward invented certainty or failed output. Allow nulls and uncertainty states where the source may not support a conclusion. Don't mistake a required field for a verified fact.

Setting a high cap as a substitute for prompt quality. Extra budget cannot repair an ambiguous category, a missing exclusion rule, an inadequate identifier or a weak acceptance test. Check the first examples and adjust the task before increasing spend.

Treating any completed response as exhaustive. Inspect the terminal stop reason and the returned count against your target. A budget-limited response can contain usable work and remain incomplete.

Checking only citations exist. Open the source. Test whether it supports the exact field and entity. A URL can be valid and irrelevant, or a page can describe the category without proving that a specific vendor qualifies.

Comparing settings on different tasks. Run equivalent inputs and hold dates, schema, exclusions and review policy stable. If the task is stochastic or the web changed, repeat enough representative trials to understand the variation.

Letting partial results feed downstream automation. Keep an explicit human checkpoint. A budget cutoff should be visible in the queue and on every handoff, not buried in logs no one opens.

Using the benchmark as a promise for your workload. Exa's launch report is an Exa report, and the independent launch coverage reviewed here did not reproduce the tests. The WANDR study itself shows that high-effort search still has coverage and evidence gaps. Your workload needs its own sample and acceptance rule.

A run should know how it can fail

Choose Actual size to read the graphic closely.

The useful boundary of Agent Ultra is not simply its highest effort. It is whether the team can explain the extra research budget, measure a result that matters and protect downstream decisions when the run stops short. Before a run, write down the schema, cost ceiling, time limit, task boundary, evidence checks and owner of the next state. Afterward, keep the stop reason, billing and verification outcomes attached to the run ID.

A capped run can be a sensible research assistant even when it is incomplete. It gives the team a candidate set to inspect and a traceable basis for deciding whether more research is worthwhile. It becomes unsafe when the workflow suppresses the limit, treats valid JSON as verified facts, or turns a large returned count into a claim of complete coverage. Partial evidence is still useful when its status is honest.

For mode selection, the companion article “Exa Agent Ultra: When Is Exhaustive AI Research Worth It?” compares the workloads that may justify Ultra over Auto or fixed effort. This guide stays on run controls and result handling. For a broader example of separating usage prices from accepted-task costs, see What Does an AI Task Cost?.

Checked for this article

Sources

  1. Exa Agent quickstartexa.ai
  2. Exa's current Ultra guideexa.ai

Keep going

All articles