Lifestyle

Cloudflare launches Auto Router public beta: AI Gateway now picks an AI model for every request automatically

On September 30, 2026, Cloudflare announced a public beta of Auto Router, an AI Gateway feature it says picks a model for each request based on how hard the task is, to cut companies' AI spending. Here is how it works, what Cloudflare's own test data shows and what it means for workers and businesses.

About 6 min read

Cloudflare launches Auto Router public beta: AI Gateway now picks an AI model for every request automatically
Image: Mokaair (Original editorial artwork)

What happened

On September 30, 2026, Cloudflare announced on its official blog that it is launching Auto Router as a public beta, meaning a test version open to users, through its AI Gateway. AI Gateway is Cloudflare's service through which an organization's AI requests pass, and it lets the organization set budgets and spending limits. According to Cloudflare, developers only need to set the model to cloudflare/auto, and Auto Router will automatically route each request to a model capable of handling that task, so end users don't have to decide which model to use themselves.

Cloudflare explained the reasoning behind the launch: in many tools such as OpenCode, Claude Code and Codex, users still have to pick a model manually, and they often end up using a model that is more powerful than the job requires. Cloudflare said early results from its internal use via OpenCode show cost savings of up to 30% compared with using only frontier models, the most capable top-tier models such as OpenAI Sol and Anthropic Claude Opus. Cloudflare also said Auto Router is free during the beta.

Cloudflare launches Auto Router public beta: AI Gateway now picks an AI model for every request automatically
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

Test data published by Cloudflare

Cloudflare used its internal general knowledge-work benchmark to compare cloudflare/auto, OpenAI GPT-6 Sol and Anthropic Claude Opus 5.5. According to Cloudflare, the test used simulated workspace tools covering everyday workflows such as email, calendar, Slack, files, travel and finance, with 97 tasks in total, and each model was sampled three times per task.

Source: internal benchmark results published on Cloudflare's official blog; not an independent test.
ModelSuccessesSuccess rate (95% confidence interval)Total costCost per success
cloudflare/auto252/29186.6% (+6.2/−6.9 percentage points)US$2.10US$0.0084
Anthropic Claude Opus 5.5281/29196.6% (+2.7/−3.8 percentage points)US$5.91US$0.0210
OpenAI GPT-6 Sol245/29184.2% (+6.5/−6.9 percentage points)US$2.64US$0.0108

Cloudflare says Auto Router costs about 80% as much as Sol and 35% as much as Opus. The same table also shows that Claude Opus 5.5 has the highest success rate of the three, and that Auto Router's advantage lies mainly in cost. Cloudflare further notes that a lower price per token (the small chunks of text that AI models process and charge by) does not necessarily mean a lower overall cost, because a seemingly cheaper model may consume more tokens to complete a task.

How Auto Router works

According to Cloudflare, the process breaks down roughly into the following steps:

  1. Building the model pool: AI Gateway excludes models that don't support the request format or execution mode, and takes into account credentials, billing settings, access control policies and spending caps; during outages it also excludes unhealthy providers or models.
  2. Classifying the request: the conversation is sent to a multi-head classification model (a single AI model that makes several judgements at once) running on Workers AI. It estimates the probability that the request belongs to each of 14 task categories and scores complexity, ambiguity, stakes and dependence on prior context on a scale of one to five.
  3. Calculating utility: Cloudflare says the router selects a model using "utility = expected quality − adaptive cost penalty"; price carries more weight for simple requests, and the cost penalty decreases as difficulty increases.
  4. Accounting for caching costs: a model keeps the conversation so far in a cache it can reread cheaply, and switching models means the new model must write it all again. So within a single turn (one round of user input) Auto Router tends to stick with the same model, while across turns it applies a switching penalty that grows with the number of tokens already in the conversation.
  5. Fallback: the router returns a ranked list, and if the provider of the top-choice model cannot serve the request, AI Gateway can switch to another eligible model.

Cloudflare says that when new models are released, no retraining is needed; their benchmark weights only need to be added to the scoring matrix.

What it means for workers and businesses

For the average office worker, the significance of this kind of tool is that in the future, your company's internal AI assistant may no longer let you decide which model to use; instead, the system will allocate one automatically based on task difficulty. Simple tasks such as sorting email might go to a smaller model, while complex work can still use a more powerful one. For businesses responsible for AI budgets, Cloudflare's pitch is to automatically cut unnecessary spending without restricting employees' access to high-end models. However, results still depend on each organization's actual workload, and for now there are only Cloudflare's self-reported figures.

What's next

  • Expand the models available to cloudflare/auto
  • Incorporate zero data retention requirements when filtering models
  • Factor in provider capacity when selecting models
  • Choose an appropriate reasoning level for each request
  • Fully support the Responses API and WebSockets
  • Explore using a structured decision model as a first-layer classifier
  • Launch cloudflare/auto-best, which pursues only the highest expected quality without weighing cost trade-offs

Frequently asked questions

What is Auto Router?

According to Cloudflare, it is an AI Gateway feature: once the model is set to cloudflare/auto, it automatically routes each request to a model capable of handling that task.

Can it really save 30%?

Cloudflare says its early internal results show savings of up to 30% compared with using only frontier models. This is a vendor-reported upper-bound figure that has not been independently verified, and actual results vary by use case.

Does automatic model selection lower quality?

In the tests Cloudflare published, Auto Router had a success rate of 86.6%, higher than GPT-6 Sol's 84.2% but lower than Claude Opus 5.5's 96.6%. Cloudflare also plans to launch cloudflare/auto-best, which ignores cost and pursues only the highest quality.

Do I have to pay for it now?

Cloudflare says Auto Router is free during the beta. Cloudflare did not explain in the post how it will be priced after the beta ends.

Does switching models waste money?

Cloudflare says Auto Router factors in cache read and write costs: within a single turn it tends to stick with the same model, and across turns it raises the switching threshold based on conversation length.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle