Skip to main content

AI in Practice

Exa Agent Ultra: When Is Exhaustive AI Research Worth It?

When is Exa Agent Ultra worth it? See what its $20 default cap and vendor-reported benchmarks mean, then test it against a completeness target.

Exa Agent Ultra cover: a field of research cards narrows through a verification gate into a smaller checked set, with the official Exa mark.
On this page
  1. What changed with Exa Agent Ultra
  2. Why the benchmark story deserves a pilot, not a purchase decision
  3. The right task is defined by the cost of an omission
  4. Compare coverage and correctness together
  5. A worked decision, without pretending it is a result
  6. How to choose between Ultra, Auto and fixed effort
  7. A practical pilot checklist
  8. The decision is a workload rule, not a winner badge

Exa Agent Ultra is Exa Agent’s highest-effort research mode. It is worth testing when your task is a broad, hard-to-complete list and the cost of missing valid records matters more than a fast answer. It is not automatically the right choice for every research query. Exa introduced Ultra on September 25, 2026 as the highest-effort mode of its Agent API. Its own benchmark report says Ultra led selected wide-research tests, but Exa adapted the benchmark harness and says some comparison results came from other vendors' published runs. There is no independent reproduction of those comparative results in the coverage reviewed here. Treat them as a reason to run a careful pilot, not proof that Ultra is the best fit for your work.

The practical question is not “Which agent won a chart?” It is “What is one missed or incorrect result worth in this workflow?” A team mapping a market for a diligence review may care about long-tail coverage. A support manager looking up one current policy may need an answer quickly and can verify it directly. Those jobs should not inherit the same effort setting just because the same API accepts both.

Research checked September 30, 2026. Prices and product behavior can change. No Exa run or comparative hands-on benchmark was performed for this article.

Key takeaways- Reserve Ultra for broad research where omitted qualifying records carry a real cost.- Treat Exa’s benchmark comparisons as vendor-reported until independently reproduced.- Compare accepted, source-verified records and review labor against a fixed or Auto baseline.

What changed with Exa Agent Ultra

Choose Actual size to read the graphic closely.

Ultra is a high-effort setting inside Exa Agent, not a separately hosted product or a new foundation model. Exa describes Agent as a service that can split a task into subtasks, research multiple domains, assemble structured output, and return grounding details. Ultra asks the service to spend more effort on broad, deep research where thoroughness is more important than latency or cost. The Agent API documentation currently describes effort: "ultra", optional outputSchema, input rows, and streamed or asynchronous runs. It also says typical complex Ultra runs take about 30 minutes, while especially challenging work can run up to three hours. The duration ceiling can be set between five minutes and three hours; there is no default duration ceiling listed in the current Ultra guide. Exa's Agent Ultra documentation is the current reference for the run controls.

The billing model is metered. Exa's Ultra page lists a default maximum cost of $20 per run. A caller can set a cost cap between $1 and $100, subject to account or server limits, and can set a maximum duration. A cap is a ceiling, not a quoted price. A run that completes early may cost less; a run that reaches its cap can stop starting new work and return what it has gathered. Exa's Agent guide describes usage components and recommends fixed-effort options when predictable per-request pricing matters. Do not compare the $20 cap with a fixed per-request fee as if both were the expected cost of every call.

There is a useful choice behind the list of effort settings. The API supports fixed efforts such as low, medium, high, and xhigh, plus auto and ultra. Fixed settings provide predictable request prices in the current docs. Auto is metered with a default $5 cap and is intended for scope that varies. Ultra has a $20 default cap and is intended for exhaustive work. Those are product recommendations from Exa, not a rule that automatically selects correctly for a particular job. Your task's schema, count target, field difficulty, review standard, budget and acceptable wait are still your decisions. The Agent API quickstart describes structured output and field-level grounding; those features are valuable only if you inspect whether the output satisfies your actual criteria.

Why the benchmark story deserves a pilot, not a purchase decision

Choose Actual size to read the graphic closely.

Exa's September 25 launch post reports results on WANDR, DeepSearchQA, WideSearch and a Company Find-All evaluation. It presents comparisons with Opus 5.5, GPT-6 Astra and Perplexity Agent, including task-level scores, task costs and a large average count of qualifying companies found. Those are consequential claims, so the evaluation details matter at least as much as the winning numbers.

Exa says the WANDR grader shares the upstream evaluation logic, but its comparison changes the contents tool, transport logic and judge model. Exa reports another vendor's result when that vendor had published a result on the same grader harness; where a result was not available, Exa ran the benchmark itself. The published article therefore combines different evidence origins. It is not a single, independent, head-to-head rerun of every competing system under an identical live configuration. The right summary is narrower: Exa reports strong results on its selected evaluation setup, while the authors disclose meaningful harness choices that limit how directly a buyer can generalize the leaderboard.

The WANDR research paper, published on arXiv in August 2026, is useful context because it explains the benchmark's task shape. It contains 500 realistic wide-and-deep data-collection tasks. A task can require discovering many entities, enriching each entity across multiple fields, and supplying cited excerpts for each record. Its evaluator fetches sources and checks both the claims and evidence. WANDR reports that research agents still have substantial headroom: even the strongest system in its paper achieved 0.363 soft F1 and 0.133 hard F1 at high effort. That paper is not an independent validation of Agent Ultra. It is a benchmark paper showing why a row count, a plausible answer or a single average score can hide missing records and incomplete evidence.

MarkTechPost's September 26 launch coverage provides an independent editorial account of Exa's announcement and reproduces the benchmark table and Exa's methodology description. It explicitly says the results remain vendor-reported and have not been independently reproduced. Independent coverage helps confirm what was announced and how the claims were presented; because it relies on Exa's reported numbers, it does not constitute a second benchmark run. That distinction matters: a reporter independently describing a vendor's result is not the same as a lab independently measuring it.

The right task is defined by the cost of an omission

Choose Actual size to read the graphic closely.

The strongest fit is a request with an open or large result set, criteria that require judgment, and an operational reason to pursue more coverage. For example, a team might want candidate vendors that satisfy several eligibility conditions, with each row supported by current public evidence. Or a research group may need a catalog of relevant papers where each record has to match a defined method and link to an accessible source. In both cases, a short answer containing a few obviously relevant names may be inadequate even if every included name is correct.

That doesn't mean any request with the word “all” belongs on Ultra. “Find all providers” is not a measurable task until you state the population, geography, date window, definition of provider, qualifying evidence and stopping rule. Exhaustiveness over the entire public web cannot usually be proven by a model. You can define a benchmarkable target: coverage of a known, independently assembled sample, or a required number of qualifying items within a clear, finite universe. When the population itself is not observable, report a bounded search outcome, not a claim of completeness.

A one-off lookup has different economics. If the reader needs a current product limit or a particular company's stated price, opening the official page, checking the relevant section and summarizing it may be enough. A high-effort multi-agent run could do more work without changing the decision. The cost of extra completeness is justified only when it lowers a meaningful omission risk, improves evidence quality or saves enough human research time to offset the run cost and review burden.

Think of the total job, not only the API invoice. The run may use credits or paid partner tools; a human may still need to resolve duplicates, inspect citations, fix schema issues and confirm the right entities. In the opposite direction, the run can save time by assembling a first pass that would otherwise require repeated searches. Neither outcome is guaranteed by effort mode. The pilot should record both compute spend and the staff time required to turn its output into an accepted research artifact.

Compare coverage and correctness together

Choose Actual size to read the graphic closely.

A task-specific evaluation needs two axes. Coverage asks how many genuinely qualifying records the system found against a target or independently built reference sample. Correctness asks how many returned records meet the criteria and whether the cited source actually supports the fields. Precision, recall and their balanced F1 score are familiar ways to summarize the trade-off when you have a defensible denominator. If you have no known universe or validated sample, avoid presenting “recall” as an observed percentage. You can still measure the number of verified accepted records, but label it as an output count rather than coverage against all possible answers.

WANDR's paper formalizes a related distinction. Its soft measures can give partial credit to records that are partly correct, while its hard measures require complete records under the benchmark's stated hierarchy. That is a valuable reminder for internal evaluation: a company name found without the required qualification evidence is not necessarily a successful company record. It may be a lead for review. Decide in advance what constitutes passing, then avoid changing the denominator or field requirements after seeing which system looks better.

A small, manually verified reference sample is often more useful than a giant table nobody checks. Take examples from different parts of the expected search space: common and obscure entities, recent and older sources, ambiguous category boundaries, multiple source types, and cases where an apparent match fails one required condition. Have a reviewer who did not prompt the model mark each record and the evidence independently. For omissions, search that sample without looking at the agent output first. Otherwise, the evaluation can inherit the model's blind spots and mistake them for the world's boundaries.

Keep identity errors visible. Similar company names, acquired brands, regional subsidiaries and products that share a parent organization can create duplicate or mismatched rows. Count duplicates according to a written entity-resolution rule, not by informal cleanup. An answer can have high apparent breadth while including the same business twice under different names. At the same time, aggressive deduplication can collapse two genuinely separate legal or product entities. Preserve a reason and source whenever a record is merged or excluded.

A worked decision, without pretending it is a result

Choose Actual size to read the graphic closely.

Imagine an operations leader needs an initial map of publicly documented vendors that support a particular workflow and satisfy four requirements. The list will be used to prioritize diligence, not to approve vendors. A human will verify any company before a purchasing decision. A first pass of ten strong candidates might be useful, but the team has repeatedly found that missing a niche provider makes the market map less useful. This scenario is hypothetical; it is not a measured Exa test.

The team should start by writing the required record: company identity, the four criteria as separate fields, a supporting URL and a short evidence excerpt for each criterion, last-checked date, and an uncertainty state. A row with an unverifiable criterion remains “needs review,” not “pass.” Then the team should establish what coverage means. Perhaps it has a seed set of 30 known providers across the four criteria and expects the research run to discover additional candidates. The seed set can test whether an agent rediscovers known examples, but it cannot establish the size of the unknown long tail. So the pilot should report two numbers separately: seed-set recovery and the number of newly found candidates that pass independent checking.

Now compare a fixed or Auto effort run with Ultra on the same frozen task. Keep the prompt, schema, exclusions, source window and review rules stable. Randomly choose which outputs receive which anonymized reviewer if practical. Record cost, runtime, number of raw rows, duplicates, rows that satisfy every criterion, rows with source support, and manual correction minutes. If different settings use materially different search scope or freshness, document it rather than claim an isolated effort comparison.

The result might justify Ultra if it produces meaningfully more verified qualifying candidates without a disproportionate increase in review labor. It might instead show that Auto reaches the operational threshold at much lower cost, or that both modes miss the same difficult category and require a better search plan. If Ultra finds more rows but its extra rows have weak support, the total is not a win. If it costs more yet moves the market map past a decision-relevant coverage threshold, that may still be a rational trade. The answer depends on the decision the map will influence, not on a benchmark headline.

How to choose between Ultra, Auto and fixed effort

Choose Actual size to read the graphic closely.

Use a fixed effort when the job is clearly bounded, its expected schema is modest and a predictable request cost is valuable. A single entity, a short list of known products, or a request that has a small number of verifiable fields may not need an open-ended research run. Start with the lowest setting that still meets the evidence and output requirements, then escalate only after measuring a failure mode that a higher effort can plausibly address.

Use Auto when the task's size or difficulty varies, but you still want the system to allocate work within a smaller default cap. Auto can be a useful baseline for list-building because the required depth can vary. Its lower default cap does not mean it will always return complete output; the same discipline around schemas, citations and partial completion applies. Use an explicit cap where the budget must be predictable.

Use Ultra for high-value broad research when you can state what “more complete” means and why missing additional records matters. Its higher default cap and longer run behavior should be a deliberate choice. Do not make Ultra your background default for every event or every item just because it is the top mode. More computation can improve the odds of finding a long tail, but neither the benchmark nor the product description makes every specific task exhaustive or correct.

If the answer needs to be produced in seconds, an asynchronous run lasting around half an hour on a complex task may be the wrong interface even if the output could be thorough. If a record can trigger a legal, compliance, hiring, medical or financial action, a model result still needs domain-appropriate human validation. “Highest effort” is not a review waiver. Keep the level of assurance proportional to the consequence and to what the source evidence establishes.

A practical pilot checklist

Choose Actual size to read the graphic closely.

Before the run, freeze the question, population, inclusion and exclusion rules, search date, output schema and accepted-record definition. Decide which fields require direct citations and what the system should return when a criterion cannot be verified. Define the maximum acceptable cost and runtime. Choose a baseline based on the current workflow rather than a straw alternative: if the team already uses Auto, compare against Auto; if it uses a human researcher, capture a timed sample of that process under the same task instructions.

During the pilot, retain the initial request and all returned structured data. Record the stop reason, actual charge, start and finish times, schema failures, source failures, duplicate handling and every human correction. A run that finishes with budget_reached or time_limit_reached may still be useful, but it is a partial result. Do not silently treat it as a complete list. The current Ultra guide says other terminal reasons can also return whatever the system found, and it distinguishes schema_satisfied, stopped, error and cancelled; interpret each according to the published meaning.

After the run, independently review a sample of positive records and likely edge cases. If possible, have someone make a blind search for omitted entries from a known list. Calculate the actual incremental cost per accepted record, including review time if it matters to the workflow. Preserve failed examples: which requirement was misunderstood, which citations did not substantiate the field, which candidate was a duplicate and what did the system stop before finishing? These cases tell you whether to adjust effort, task scope, schema, source strategy or review rules.

Run at least a second representative task before setting a recurring policy. One market map can be unusually easy or difficult. A robust decision should survive different topic areas, evidence types and source freshness. If the team cannot construct a denominator for completeness, describe a bounded acceptance threshold and keep the label “not a proof of total completeness.” Re-evaluate when the API, its pricing, your workflow or the cost of omissions changes.

The decision is a workload rule, not a winner badge

Choose Actual size to read the graphic closely.

Exa Agent Ultra gives Exa Agent a mode explicitly aimed at high-effort, long-running research. Exa's own launch comparisons are interesting, and MarkTechPost's coverage reports the announcement and its limits. Neither source demonstrates that your task will be cheaper, more complete or more accurate than your current process. The WANDR research helps explain how hard wide-and-deep work is: getting names is only one part; evidence completeness and qualification matter for each row.

My practical recommendation is to reserve Ultra for a task where the team can define the required record and explain the cost of omissions. Compare it with Auto or the existing workflow on a frozen representative sample. Keep run caps and partial-result states explicit, review sources independently, and judge the output by accepted records and remaining staff work. If Ultra does not cross a useful threshold, do not pay for effort that does not improve your decision. If it does, document the workload boundary and rerun the comparison when that boundary or the product changes.

The companion guide, “How to Set Safe Budgets for Exa Agent Ultra Research,” focuses on choosing a schema and safely handling cutoff results. For related Rise guidance on comparing AI costs, see What Does an AI Task Cost?, which shows why cost per request and cost per accepted outcome answer different questions. Its GPT-6.1 Sol cost-accounting guide is a separate example of attaching retries and review to the work being accepted; Exa Agent follows its own metered rate schedule.

Checked for this article

Sources

  1. Exa's Agent Ultra documentationexa.ai
  2. Agent API quickstartexa.ai
  3. Exa's September 25 launch postexa.ai
  4. WANDR research paperarxiv.org
  5. MarkTechPost's September 26 launch coveragemarktechpost.com

Keep going

All articles