Cut your AI spend with AI Gateway’s Auto Router

Post Syndicated from Ming Lu original https://blog.cloudflare.com/auto-router/

From our conversations with companies at every stage of their AI adoption journey, we've seen some common patterns. First, there is an exploration period as you bring on every new tool, dole out API keys freely, and let the tokens flow. Then, you converge on the canonical tools for your organization for agentic coding, for non-technical workflows, for running and deploying agents. As companies formalize their AI adoption, they want to manage and oversee token spend for users, but budgets and rules only go so far. The best savings are the ones users never notice.

Today, we are releasing Cloudflare's Auto Router in public beta, available through AI Gateway. Set your model to cloudflare/auto and the Auto Router will automatically route each request to a model that is capable enough for the task, without requiring an end user to think about model selection. Our early results using the Auto Router internally through our OpenCode harness show a cost savings of up to 30% when compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.

Why we built this

From our own experience tracking AI spend at Cloudflare, we’ve learned managing costs requires a multipronged approach. Previously, we talked about how to set budgets and limits around AI spend, and how to see who is spending across your organization by linking employees to their AI usage.

In many harnesses, including OpenCode, Claude Code, and Codex, individual users still select models manually. Of course, not all tasks are created equal, and often individuals end up using models that are overkill for their work. For example, you don't need Opus-level intelligence if you're looking to summarize an email or chat threads. However, you wouldn't want to block that model completely from your security engineering team.

Our goal is for AI Gateway to be the control plane for organizations deploying AI internally. Because every request from every user, agent, and tool already flows through it, AI Gateway is in a unique position to do more than observe and enforce. Budgets, spend limits, and identity-aware analytics give organizations visibility and guardrails, but they still rely on individuals to make cost-conscious choices request by request. The next step is for the gateway itself to make intelligent decisions on a user's behalf: sending each request to a model that is capable enough for the task. That way, organizations reduce spend automatically, while users keep access to the most capable models when their work actually needs them.

The results

We use Auto Router internally at Cloudflare within our OpenCode deployment and within Cloudflare OS, our custom agent harness. In our internal usage, we’ve seen results comparable with frontier models for coding tasks.

Auto Router does best when used across a wide range of knowledge-work tasks, like those typically found in a large organization with work spanning both technical and non-technical teams. We evaluated cloudflare/auto against OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5 on our internal general knowledge work benchmark. The benchmark uses simulated workspace tools and covers common day-to-day workflows across email, calendars, Slack, files, travel and finance. Each task requires the model to use these tools to produce a verifiable answer or complete an action.

Model

Successful Trials

Success Rate

Total Cost

Cost per success

cloudflare/auto

252/291

86.6% (+6.2/−6.9 pp)

$2.10

$0.0084

Anthropic Claude Opus 5.5

281/291

96.6% (+2.7/−3.8 pp)

$5.91

$0.0210

OpenAI GPT-6 Sol

245/291

84.2% (+6.5/−6.9 pp)

$2.64

$0.0108

97 tasks with three samples per model per task. Parenthetical values show 95% confidence intervals estimated from 10,000 task-level bootstrap resamples, preserving all three repetitions within each task. “pp” indicates percentage points.

Our Auto Router delivered similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus. While that may initially seem surprising, one way to frame the problem a model router solves is through the “jagged frontier” across models. The ability to solve a problem often exists somewhere in this portfolio of models; the router’s job is to choose the right model for each task while balancing quality and price. Savings come from not paying frontier rates for non-frontier work, and they grow with how much of that work you have.

Another insight is that lower token prices do not always produce lower-cost outcomes. A model that looks cheaper on paper may end up using disproportionately more tokens to solve a problem. A router should minimize predicted trajectory cost, not just load-balance by dollars per million tokens. 

This is already useful today, but it’s only the beginning of what the Auto Router can learn from Cloudflare’s position in the inference path.  

How it works

When you send a request to cloudflare/auto, AI Gateway first builds the pool of models that can actually serve it. It filters out models that do not support the request format or execution mode, and accounts for the credentials, billing configuration, access control policies, and spend limits attached to the gateway. It will also filter out unhealthy upstream providers or models during downtime and automatically bring them back into the pool after an outage.

For the remaining candidates, the router looks at a compact view of the conversation. It considers the most recent messages, prioritizing the newest turns. The conversation is then sent to a multi-head classification model running on Workers AI and deployed on GPUs across our edge network. The classifier produces two sets of signals. First, it assigns probabilities across 14 task categories (like coding, planning, research, data analysis). It then rates the request across four dimensions on a scale from one to five: complexity, ambiguity, stakes, and dependence on earlier context.

A separate scoring matrix combines those signals with model benchmark results to estimate how well each model fits the request. To calibrate the scoring matrix, we defined the preferred model for a set of example task and difficulty profiles, then adjusted the weights to produce those choices.

Finally, the router combines expected quality with each model's input and output token prices. On straightforward requests, price carries more weight, so a smaller model can win when it is capable enough. As difficulty rises, the cost penalty falls and stronger models have more room to win. In simplified terms, cloudflare/auto selects the model with the highest utility as defined by:

For long agentic sessions like debugging or coding, cost is less driven by the model’s list price than by the cost of cache reads, which grows with session length. Switching models throws the cache away and forces a new model to write the whole context again. This can be worth it, as a model with a cheaper cache-read and cache-write prices can pay back the rewrite quickly.

Rather than completely avoiding model switching, the Auto Router accounts for the cost of cache reads and writes. Within a turn (one user input loop), the cache is hot and switching rarely pays off, so it’s better to keep using the same model. Across turns, the Auto Router applies a switching penalty that grows with the number of tokens already in context. A model that still holds a live cache for the session is priced at its cheaper cache-read rate. Every other candidate is priced at the full cost of rewriting the context, so the deeper the conversation, the more a switch has to earn back, through higher quality results that use fewer tokens overall or cheaper cache rereads. Switching models has another cost: most models can't read another model's reasoning tokens, so a model switch that drops reasoning tokens means that the new model may have to redo it at output prices. In the future, we want to account for this by having the router prefer to stay within the same model family when it switches.

From there, the router returns a ranked list. AI Gateway attempts the winner first and can move to another eligible model if that provider cannot serve the request.

This overall design has several benefits. The two-stage architecture (task and dimensions classifier to scoring matrix) means that routing decisions are legible because you can inspect each task’s predicted category and complexity to see how it translated into the model choice. Adjusting the router when a new model is released also does not require retraining — we only add its benchmark-derived weights to the scoring matrix. The same classifier can also support different routing profiles. For example, in addition to cloudflare/auto, we plan to release other routers in the future, including cloudflare/auto-best, which uses the same classification and model pool, but selects the highest expected quality without applying the cost tradeoff.

What's next

Our release today is only the starting point, and we’re continuing to invest in research and new routing strategies. In the near term, we want to:

  • Expand the models offered through cloudflare/auto
  • Include zero-data-retention requirements when filtering models
  • Account for provider capacity when selecting models
  • Select the appropriate reasoning or thinking level for each request
  • Add full support for the Responses API and WebSockets
  • Explore structured decision models as a first-pass classifier

The Auto Router is free while in beta. Read more in our developer documentation.

Acknowledgements: This project was also made possible by the efforts of Mats Dodd, Sam Scott, Oliver Yu, and Jeff Rafter.