Lifestyle
Cloudflare launches Auto Router public beta: AI Gateway now picks an AI model for every request automatically
On September 30, 2026, Cloudflare announced a public beta of Auto Router, an AI Gateway feature it says picks a model for each request based on how hard the task is, to cut companies' AI spending. Here is how it works, what Cloudflare's own test data shows and what it means for workers and businesses.
About 6 min read

What happened
On September 30, 2026, Cloudflare announced on its official blog that it is launching Auto Router as a public beta, meaning a test version open to users, through its AI Gateway. AI Gateway is Cloudflare's service through which an organization's AI requests pass, and it lets the organization set budgets and spending limits. According to Cloudflare, developers only need to set the model to cloudflare/auto, and Auto Router will automatically route each request to a model capable of handling that task, so end users don't have to decide which model to use themselves.
Cloudflare explained the reasoning behind the launch: in many tools such as OpenCode, Claude Code and Codex, users still have to pick a model manually, and they often end up using a model that is more powerful than the job requires. Cloudflare said early results from its internal use via OpenCode show cost savings of up to 30% compared with using only frontier models, the most capable top-tier models such as OpenAI Sol and Anthropic Claude Opus. Cloudflare also said Auto Router is free during the beta.
Read the full description
Sources are collected, independently checked, then reviewed by Jev.
Test data published by Cloudflare
Cloudflare used its internal general knowledge-work benchmark to compare cloudflare/auto, OpenAI GPT-6 Sol and Anthropic Claude Opus 5.5. According to Cloudflare, the test used simulated workspace tools covering everyday workflows such as email, calendar, Slack, files, travel and finance, with 97 tasks in total, and each model was sampled three times per task.
| Model | Successes | Success rate (95% confidence interval) | Total cost | Cost per success |
|---|---|---|---|---|
| cloudflare/auto | 252/291 | 86.6% (+6.2/−6.9 percentage points) | US$2.10 | US$0.0084 |
| Anthropic Claude Opus 5.5 | 281/291 | 96.6% (+2.7/−3.8 percentage points) | US$5.91 | US$0.0210 |
| OpenAI GPT-6 Sol | 245/291 | 84.2% (+6.5/−6.9 percentage points) | US$2.64 | US$0.0108 |
Cloudflare says Auto Router costs about 80% as much as Sol and 35% as much as Opus. The same table also shows that Claude Opus 5.5 has the highest success rate of the three, and that Auto Router's advantage lies mainly in cost. Cloudflare further notes that a lower price per token (the small chunks of text that AI models process and charge by) does not necessarily mean a lower overall cost, because a seemingly cheaper model may consume more tokens to complete a task.
How Auto Router works
According to Cloudflare, the process breaks down roughly into the following steps:
- Building the model pool: AI Gateway excludes models that don't support the request format or execution mode, and takes into account credentials, billing settings, access control policies and spending caps; during outages it also excludes unhealthy providers or models.
- Classifying the request: the conversation is sent to a multi-head classification model (a single AI model that makes several judgements at once) running on Workers AI. It estimates the probability that the request belongs to each of 14 task categories and scores complexity, ambiguity, stakes and dependence on prior context on a scale of one to five.
- Calculating utility: Cloudflare says the router selects a model using "utility = expected quality − adaptive cost penalty"; price carries more weight for simple requests, and the cost penalty decreases as difficulty increases.
- Accounting for caching costs: a model keeps the conversation so far in a cache it can reread cheaply, and switching models means the new model must write it all again. So within a single turn (one round of user input) Auto Router tends to stick with the same model, while across turns it applies a switching penalty that grows with the number of tokens already in the conversation.
- Fallback: the router returns a ranked list, and if the provider of the top-choice model cannot serve the request, AI Gateway can switch to another eligible model.
Cloudflare says that when new models are released, no retraining is needed; their benchmark weights only need to be added to the scoring matrix.
What it means for workers and businesses
For the average office worker, the significance of this kind of tool is that in the future, your company's internal AI assistant may no longer let you decide which model to use; instead, the system will allocate one automatically based on task difficulty. Simple tasks such as sorting email might go to a smaller model, while complex work can still use a more powerful one. For businesses responsible for AI budgets, Cloudflare's pitch is to automatically cut unnecessary spending without restricting employees' access to high-end models. However, results still depend on each organization's actual workload, and for now there are only Cloudflare's self-reported figures.
What's next
- Expand the models available to cloudflare/auto
- Incorporate zero data retention requirements when filtering models
- Factor in provider capacity when selecting models
- Choose an appropriate reasoning level for each request
- Fully support the Responses API and WebSockets
- Explore using a structured decision model as a first-layer classifier
- Launch cloudflare/auto-best, which pursues only the highest expected quality without weighing cost trade-offs
Frequently asked questions
What is Auto Router?
According to Cloudflare, it is an AI Gateway feature: once the model is set to cloudflare/auto, it automatically routes each request to a model capable of handling that task.
Can it really save 30%?
Cloudflare says its early internal results show savings of up to 30% compared with using only frontier models. This is a vendor-reported upper-bound figure that has not been independently verified, and actual results vary by use case.
Does automatic model selection lower quality?
In the tests Cloudflare published, Auto Router had a success rate of 86.6%, higher than GPT-6 Sol's 84.2% but lower than Claude Opus 5.5's 96.6%. Cloudflare also plans to launch cloudflare/auto-best, which ignores cost and pursues only the highest quality.
Do I have to pay for it now?
Cloudflare says Auto Router is free during the beta. Cloudflare did not explain in the post how it will be priced after the beta ends.
Does switching models waste money?
Cloudflare says Auto Router factors in cache read and write costs: within a single turn it tends to stick with the same model, and across turns it raises the switching threshold based on conversation length.
Browse the latest news in this topic
Lifestyle
Cloudflare open-sources Streamline: a demo of using its cloud services to add graphics to live streams and burn subtitles into videos
On October 2, 2026, Cloudflare launched and open-sourced Streamline, a developer playground showing how developers can combine Stream, Workers, Containers and Durable Objects to build their own video processing pipelines, such as adding graphics to live streams in real time or adding subtitles to videos. This article explains what it is, how it works, its limitations, and what it means for viewers and developers.
Lifestyle
Google unveils Gemini 4 Argon: cyber defenders get it first, everyone else still has to wait
On September 30, 2026, Google announced Gemini 4 Argon, which it calls its new frontier (most advanced) AI model. For now it is available only to trusted cyber defenders through the Fairwind Program. Here is what Google says the model can do, what it will cost developers, how Google says it is managing the risks, and what it means for everyday users. All figures come from Google itself.
Lifestyle
NVIDIA: CoreWeave Begins Offering Vera Rubin NVL72, With Cognition as First Production Customer
According to the NVIDIA blog, AI cloud provider CoreWeave now offers NVIDIA's next-generation Vera Rubin NVL72 systems, plans to offer the Vera CPU and launched CoreWeave Forge. This matters mainly to companies building AI agents, and could eventually mean faster AI tools for everyday users.
Lifestyle
Google Cloud makes Spanner Omni generally available: its Spanner database can now run in companies' own data centers and on other clouds
Google Cloud says Spanner Omni, a version of its Spanner database that businesses run themselves, is now ready for real-world use in their own data centers, on other clouds or on a laptop. This matters to organizations that want Spanner outside Google Cloud, but they take on the running of it. Here are the features, licences and trade-offs Google Cloud describes.
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family