Skip to main content

Automation and Agents

How to Pilot Cloudflare Auto Router With Cost and Quality Guardrails

Plan a bounded Auto Router test with eligible models, privacy checks, spend limits, quality sampling, session affinity, stop rules, and rollback.

A Cloudflare-branded pilot route filters eligible providers, checks privacy and spend controls, then reaches human review.
On this page
  1. Start with a pilot contract, not a traffic switch
  2. Freeze which models may receive the request
  3. Map the data path before testing sensitive requests
  4. Put a hard edge around spend and fallback
  5. Log a decision without logging everything
  6. Keep multi-turn sessions coherent
  7. Resolve interface and compatibility questions up front
  8. Compare in stages and set stop rules
  9. What to take to the review meeting
  10. Frequently asked questions
  11. Can Auto Router be used safely with confidential data?
  12. How should I limit Auto Router spending?
  13. Does Auto Router support OpenAI Responses API?
  14. Does Cloudflare Dynamic Routing replace Auto Router?
  15. Sources

--- title: "How to Pilot Cloudflare Auto Router With Cost and Quality Guardrails" blogTitle: "How to Pilot Cloudflare Auto Router With Cost and Quality Guardrails" metaTitle: "How to Pilot Cloudflare Auto Router With Cost and Quality Guardrails" metaDescription: "Plan a bounded Auto Router test with eligible models, privacy checks, spend limits, quality sampling, session affinity, stop rules, and rollback." description: "Plan a bounded Auto Router test with eligible models, privacy checks, spend limits, quality sampling, session affinity, stop rules, and rollback." category: "Automation and Agents" author: "Demetri Panici" excerpt: "A model router changes who handles each request. Here is a practical pilot plan for testing Cloudflare Auto Router without losing track of data, quality, spend or rollback." hero: ../assets/second-cover-hero.png heroAlt: "A Cloudflare-branded pilot route filters eligible providers, checks privacy and spend controls, then reaches human review." status: private_draft ---

Cloudflare AI Gateway Auto Router can choose a different eligible model for each request, so a useful pilot has to control more than the model name. Set the pilot boundaries before any real work goes through cloudflare/auto: which requests and providers are in scope, what data those providers may receive, the test budget, the pass rule, and who can stop it. Auto Router makes routing choices; your team sets the policy around them.

Cloudflare announced Auto Router as a public beta on September 30, 2026. Its docs describe a candidate pool, task classification, quality-and-cost scoring, model selection and fallback. They also expose allowed-model and allowed-provider headers, session-affinity controls and routing-decision identifiers. Those are the ingredients of an observable test, not a complete governance program.

This guide assumes you have already decided that a routing test is worth running for one workflow. Whether Cloudflare's reported savings justify a test at all is a separate question. The job here is narrower: assemble the available controls into an experiment you can observe, stop and reverse. What follows is a plan, not an account of a pilot I configured or ran, and the running example is illustrative. Cloudflare's own benchmark supplies a reason to investigate, but its results do not predict this pilot's costs or quality.

A pilot charter fixes the workflow, owner, evaluation sample, acceptance rule and stop boundary before routing begins. View image detail

Choose Actual size to read the graphic closely.

Start with a pilot contract, not a traffic switch

A pilot contract names one workflow that is narrow enough to evaluate. “Our company uses AI” is not a pilot. “Classify incoming support tickets into these approved categories, cite the text supporting the label, and send uncertain cases to a person” is much closer. It names a request type, a measurable output and an escalation path. This guide uses that ticket classifier as its running example. Another workflow needs its own contract, because its data, errors and review burden differ.

The contract should answer five questions. Who owns the test and who can stop it? Which users, application path and data class are included? What fixed route is the baseline? What makes a result acceptable? What spending or quality limit ends the experiment? If you cannot write those answers concretely enough to put them in a log or a review checklist, the team is not ready to expose live work. And if nobody has yet asked whether the task should be automated at all, Rise's work-worth-doing test comes first.

Choose a sample that resembles the work you expect to route. Include ordinary cases and the harder cases that drive cost or failure, rather than only examples that make the feature look good. For the ticket classifier, that means routine requests alongside the ambiguous or multi-issue tickets most likely to be mislabeled. Keep test and evaluation data in the environment approved for that data. A shadow comparison, where the routed output cannot trigger an action, reduces early operational risk. If your integration cannot run in shadow mode, a small opt-in cohort with human review can still limit exposure.

Write the success rubric before you start. It may require a correct label, a traceable source, a valid structured response, no invented values, and an escalation when confidence is inadequate. Keep each criterion tied to the job. A general score such as “the answer seems fine” tends to drift after a team sees which model answered, so hide the model name from reviewers where you can. Document borderline cases so reviewers apply the same standard to baseline and candidate output.

Freeze which models may receive the request

Auto Router first builds an eligible pool. Cloudflare says eligibility can depend on supported request formats and inputs, provider credentials, billing settings, access controls, spend limits and upstream health. Its managed default pool can change over time. A pilot should therefore save the actual pool and account conditions observed at its start, then record any changes while the test is running.

Where the request or policy needs tighter bounds, Cloudflare documents cf-aig-allowed-models and cf-aig-allowed-providers. The allowed-models header replaces the default model list, and its entries can include additional models. You still need to verify exact model IDs, account access and provider behavior. A wildcard or a provider name can cover more than a single model identifier, so read the current docs and test the allowed list with a harmless request before production traffic. The docs say AI Gateway returns HTTP 400 when a cf-aig-allowed-models entry matches no model. They do not specify what happens if every listed model becomes ineligible at runtime, so test that condition before relying on it.

The eligible pool removes models that fail format, input, account, policy or health requirements before the router ranks remaining candidates. View image detail

Choose Actual size to read the graphic closely.

Keep two questions apart: “Can the model handle this request?” and “Is this provider permitted for this data and purpose?” The first is compatibility. The second belongs to the organization's policy and contracts. The router's eligibility logic can filter on the conditions you configure, but it cannot decide whether your legal team, a customer agreement or an internal rule permits a particular prompt to reach a particular provider.

Candidate-pool control also protects your results. If the list changes during a pilot, a shift in cost or success might come from a new model, a different provider condition or the router's classification, and you will not be able to tell which. Freeze the pool for the evaluation when practical. If you update it deliberately, record a new version and analyze it as a separate phase rather than blending the results.

Map the data path before testing sensitive requests

A request can contain more than the latest user message. In a coding or agent session it may include system instructions, earlier turns, retrieved records, tool outputs and customer material. Even the ticket classifier forwards whatever a customer typed, including any names or account details in the ticket body. Cloudflare's launch post says the Auto Router classifier receives a compact view of the conversation, and eligible model providers then handle the request. The exact data path and each provider's terms matter for the kind of work you send.

The launch post lists adding zero-data-retention requirements to candidate filtering among Cloudflare's future work. So a team should not assume the router currently enforces a retention requirement just because a model is in the pool or a gateway sits in the path. Check the current product behavior and each candidate provider's applicable terms. Ask the people responsible for privacy, security and customer commitments to approve the actual path, not a generalized diagram of AI Gateway.

A privacy checkpoint sits before classification and candidate providers, with a hold when any recipient fails the data policy. View image detail

Choose Actual size to read the graphic closely.

If the prompt contains sensitive information, decide whether to remove or tokenize it, use a provider-specific route with known terms, or exclude the workflow from Auto Router. Redaction can change what a task means, so test it rather than treating it as automatically safe. Avoid keeping full prompts and outputs in evaluation logs unless policy allows it and the data is needed. A stable task ID and outcome label can often support measurement without copying source content into a spreadsheet.

A provider's general privacy statement does not settle the governance question. Verify the specific account setup, contractual terms, logging configuration, retention settings and jurisdictions relevant to the workflow. If any part remains unclear, keep that data out of the experiment. A cost comparison is worthless if the work was never allowed on that route.

Put a hard edge around spend and fallback

Set a maximum budget for the experiment before it starts, and know which control enforces it. Cloudflare's AI Gateway docs describe spend limits that can be scoped by provider, model or custom metadata, and dynamic routing includes budget and rate-limit nodes with configured fallback paths. Confirm which of these applies to your exact Auto Router integration and account. Treat a capability of the broader gateway as an Auto Router option only once current documentation or observed account behavior confirms it.

Cloudflare's spend-limit documentation describes a cumulative estimated-cost budget over a rolling or fixed time window, scoped by model, provider or request metadata. It does not document a per-request cap. Cloudflare says enforcement is eventually consistent: because a request's cost is recorded after it completes, concurrent calls can briefly push the total past the configured window budget. An enforced over-budget rule returns HTTP 429 until that window resets. A separate rate-limit rule controls request volume, not dollars.

Auto Router eligibility also accounts for spend limits, but Cloudflare does not explain how eligibility filtering interacts with the spend-limit response when a candidate reaches its limit. Test that case, including what happens if all candidates are affected. If you need a per-request ceiling, implement and test that guard in the application as a separate control.

Decide what should happen before the pilot starts. An enforced spend-limit rule can block a request with HTTP 429. Cloudflare separately documents a configured Dynamic Route that can send a request to its fallback model rather than block it. Treat these as distinct configurations: the fallback is not a step that happens after an Auto Router 429. Your runbook could instead pause the workflow or send the request to a person; those are operational choices you design, not native spend-limit actions. Test the exact configured response before exposing live work.

The combined Auto Router no-candidate outcome is undocumented. Separately, a spend-limit configuration can block with HTTP 429 or use a configured Dynamic Route fallback instead. View image detail

Choose Actual size to read the graphic closely.

Rate limits protect throughput and keep a single user or process from consuming the entire pilot. If departments have different workloads, a per-team ceiling may suit them better than one shared gateway cap. If metadata such as project, user or workflow scopes a limit, confirm that the values are actually attached to requests and that the policy catches missing or malformed metadata. An absent tag should not quietly turn a bounded route into an unbounded one.

Measure what upstream providers bill, not just the router's free beta status or router fee. Include input and output tokens, cache reads or writes where charged, tool requests, retries and the cost of any replacement attempt. For an agent, the first selected model is only one step in the transaction. If an unsuccessful response makes a tool run again or a reviewer request a rewrite, those costs belong to the task. In the ticket example, a mislabeled ticket that a person must re-route is part of the cost of that request. The Rise guide to measuring an AI task's full cost offers a complementary way to set the denominator. The article on why lower model prices do not automatically make a workflow cheaper covers the same denominator from the provider-pricing side.

Log a decision without logging everything

Your test record should let a reviewer answer a short set of questions about any request. Which approved task was this, what route served it, and which model did the router choose? Did fallback occur, what did the attempt cost, and did its result pass the rubric? Cloudflare's docs describe response headers for routed model, routing reason, routing decision ID and request ID. The documented reason values are cost_optimal_within_pool, forced_by_candidate_pool, pinned_by_turn, fallback_candidate_unavailable, fallback_key_not_in_candidates, fallback_router_error, fallback_router_timeout, and fallback_unsupported_input.

A compact request log records task ID, candidate pool, selected model and reason, decision trace, attempt cost, and reviewer outcome. View image detail

Choose Actual size to read the graphic closely.

A workable record has four groups of fields:

  • Identity: a pseudonymous task ID, experiment version, timestamp, workflow class, candidate-pool version, and session or turn ID.
  • Route: the routed model and routing reason, request and decision IDs, and any error or fallback status.
  • Cost: token and tool charges, plus latency if it matters to the job.
  • Outcome: pass or fail against the rubric, failure category and reviewer time.

Use only fields you can explain and retain. For the ticket classifier, the task ID can be the ticket's internal reference rather than its text. If exact request bodies are needed for a quality audit, handle them under the same access, redaction and retention policy as the original work.

Logging has its own failure modes. A missing decision header, an incomplete billing record or a lost reviewer outcome can make the final comparison unusable. Check that the fields arrive for successful calls, provider errors, fallbacks and rate-limit events. Keep a small set of known test cases and reconcile their gateway records with the reporting view. A dashboard showing a model label helps, but it is not evidence that the same task, cost and outcome are joined correctly.

Cloudflare's product documentation and analytics may change during beta. Record the exact docs version, request format and observed header names in the pilot record. Do not assume a log field is stable because one test returned it. If a critical data point is missing, treat it as a measured limitation or a stop condition rather than filling it in from memory.

Keep multi-turn sessions coherent

Choosing a model on every individual request can create avoidable context costs. Cloudflare says Auto Router supports session and turn identifiers, so it can keep a model through a turn and reason about switching between turns. Its launch description notes that changing models may require rewriting cached context, and a new model cannot necessarily use another model's reasoning tokens. A cheaper candidate can therefore raise cost if it forces a long conversation to be rebuilt again and again.

The ticket classifier is mostly single-turn, so this section matters most when a pilot covers a coding or agent workflow. For a multi-turn workflow, define what counts as a conversation and what counts as a turn. Use the documented session ID consistently for requests that belong together. Track whether routing stays pinned through tool calls and when a new user turn allows a switch. If the client application is responsible for sending session IDs, verify that it does. Without an ID, Cloudflare's docs say requests may select different models and lose prompt-cache reuse.

Within a user turn, a model remains pinned through tool calls; a switch at the next turn is evaluated against cache and context costs. View image detail

Choose Actual size to read the graphic closely.

Do not turn session affinity off just to increase switching. Compare the real interaction pattern with and without affinity only if both modes are permitted and meaningful for your use case. Keep the same acceptance rubric and log the model and cache-related costs for each. A one-shot classification request does not tell you how routing behaves in a coding session with many tool calls and a large context window.

Resolve interface and compatibility questions up front

Cloudflare's October 2 Auto Router docs say it supports Chat Completions and Responses API formats and does not yet support WebSockets. The September 30 launch post lists “Add full support for the Responses API and WebSockets” as near-term work. The two sources' wording conflicts, and the sources reviewed for this guide do not explain whether the blog refers to a particular Responses feature while the docs describe base-format compatibility.

If Chat Completions covers the workflow, start there and avoid the open question. If your application relies on the Responses API, write down the precise operations you need and test them against the current endpoint before the pilot. Do not read a general “supported” statement as a promise about streaming, tool cycles, stored response IDs or every parameter your app uses. If the feature path cannot be verified, use a compatible alternative or hold this workflow until Cloudflare clarifies. Based on the documentation reviewed, WebSockets stay out of scope.

A candidate can also be filtered out because its supported content types do not match the request. Test the actual mix of text, images and tool schemas the workflow will send. Write the pool and request format into the pilot record so that a later model or endpoint change cannot quietly invalidate the comparison.

Compare in stages and set stop rules

A controlled pilot can move through three stages. First, check setup with harmless sample requests: confirm authentication, allowed candidates, required headers, usage tracking, fallback and budget controls. Next, replay a representative set through both the baseline and the candidate path without letting automatic actions reach customers or change live records. Finally, if both the technical and quality gates pass, expose a small opt-in cohort with human review and a clear path back to the fixed route. After those three exposure stages, use a fourth decision gate: expand only if the pre-set criteria pass, otherwise restore the approved baseline. The fourth step is a decision, not another traffic stage.

A staged ramp moves from configuration checks to offline replay to a small reviewed cohort; every stage has a rollback route. View image detail

Choose Actual size to read the graphic closely.

The right cohort size and time window depend on task volume and risk. A small fixed number of examples can expose a broken integration but cannot establish a reliable success rate for a varied workload. Continue until the sample includes the cases that matter and the outcome is stable enough to support the decision. If the workflow is low-volume, report that uncertainty instead of treating the first week as conclusive.

Set the pass criteria before you look at results. They could include a minimum task-success rate relative to the baseline, a maximum cost per accepted result, a maximum review time, and a zero-tolerance list for certain failure categories. Do not copy Cloudflare's internal result as your threshold. The bar belongs to the value and consequences of your workflow.

Stop or roll back when any of these occurs:

  • a data-policy condition is violated;
  • spending crosses the cap;
  • routing errors become frequent;
  • the candidate pool changes without review;
  • required log fields disappear;
  • the task-success floor is missed;
  • human corrections outweigh the benefit.

Keep a simple runbook that names the person authorized to stop traffic, the fixed model or human fallback, the steps to restore it, and how to confirm the old route is active. If Cloudflare dynamic route versioning is part of your design, verify the configuration and rollback behavior in your account. The docs describe versioned routes and rollback for Dynamic Routing, but that does not prove every Auto Router setup inherits the same control.

The stop/go matrix checks quality, privacy, spend and auditability separately; any red line returns the workflow to its approved baseline. View image detail

Choose Actual size to read the graphic closely.

What to take to the review meeting

At the end of the experiment, bring a compact record rather than a single savings percentage. Include the frozen task set or its approved reference, the pass rule, the candidate pool, account and API format, cost breakdown, task success, review effort, fallback events, privacy approval and stop-rule results. Separate what was measured from what was inferred. If the evidence covers only a narrow class of requests, keep the recommendation equally narrow.

A successful pilot may justify another workflow-specific test. It does not establish a company-wide default. A disappointing pilot can still be useful if it shows that one task class needs a stronger model, that the prompt or validation needs work, or that a shared gateway cannot meet the required data boundary. For sensitive or high-stakes work, and for any workflow you cannot measure well enough to see a failure, the fixed route remains the right answer until that changes. Route only when the policy and the evidence point in the same direction.

Frequently asked questions

Can Auto Router be used safely with confidential data?

The reviewed Cloudflare announcement lists zero-data-retention filtering as future work. Check the current model pool, provider terms, account settings and your organization's policy before sending confidential data. If any candidate is not approved, exclude that workflow or use a fixed approved route.

How should I limit Auto Router spending?

Configure a cumulative spend limit for a rolling or fixed window, scoped to the provider, model or request metadata you need, and use an application guard if you need a per-request ceiling. When an enforced spend-limit rule is over budget, the spend-limit docs say it returns HTTP 429. Auto Router also uses spend limits in eligibility, but the docs do not specify the combined no-candidate behavior. A cheaper-model fallback requires a configured Dynamic Route, and pausing the workflow or handing it to a person belongs in your runbook. Count upstream model, tool, retry and review costs in the total.

Does Auto Router support OpenAI Responses API?

The updated docs and launch post conflict: the docs list Responses API support, while the launch post lists full support as future work. Test the exact request features you need or wait for clarification. The docs reviewed say WebSockets are not supported.

Does Cloudflare Dynamic Routing replace Auto Router?

They are related gateway features with different controls. The Dynamic Routing documentation describes versioned flows, conditions, quotas and rollback. Auto Router handles model selection for a request. Confirm feature availability and configuration boundaries instead of assuming a dynamic-route capability automatically applies to Auto Router.

Sources

Checked for this article

Sources

  1. Cloudflare, "Cut your AI spend with AI Gateway Auto Router"Cloudflare
  2. Cloudflare Docs, "Auto Router documentation"Cloudflare Docs
  3. Cloudflare Docs, "Dynamic routing"Cloudflare Docs
  4. Cloudflare Docs, "Spend limits"Cloudflare Docs

Keep going

All articles