Lifestyle
Google Cloud says its new GKE Agent Sandbox starts AI training sandboxes 10 to 45 times faster
Google Cloud has released a version of its GKE Agent Sandbox built for training AI agents, plus a Python toolkit and ready-made integrations. It says the sandboxes start 10 to 45 times faster, so costly GPUs spend less time sitting idle. The main audience is AI labs and startups. The speed figures come only from Google Cloud's own tests.
About 7 min read

What Google Cloud announced
Google Cloud made the announcement in a blog post dated September 30, 2026. It covers three things: a version of GKE Agent Sandbox optimized for reinforcement learning, the Agent Sandbox RL orchestration SDK (a software toolkit for managing many sandboxes at once), and native integrations with common RL "gyms" and tools, which are ready-made training environments. Google Cloud says all three are now generally available. The post was written by product manager Tinsley Shi and staff software engineer and technical lead Tomer Glottmann.
Google Cloud describes the problem this way. When an AI agent is trained with RL, a model running on GPUs (costly AI chips) produces actions such as pieces of code. Those actions are then run in a sandbox, an isolated computer environment on ordinary CPUs, so the result can be scored as feedback. At large scale, the GPUs sit idle while they wait. They wait for sandboxes to cold start (boot from scratch), for many multi-gigabyte container images (packaged copies of the software each task needs) to download, and for scheduling queues to clear.
Read the full description
Sources are collected, independently checked, then reviewed by Jev.
GKE is Google Cloud's service for running Kubernetes, a widely used open system for running software across many machines. Google Cloud describes GKE Agent Sandbox as an open Kubernetes building block for running agents securely. It includes a built-in SandboxWarmPool, a "warm pool" of environments started in advance so they are ready to use. It also supports snapshots, pause and resume.
- Orchestration SDK: Google Cloud says it offers an asynchronous Python API, so researchers do not have to write Kubernetes configuration files (YAML).
- Tool integrations: Google Cloud lists Gymnasium, NVIDIA NeMo Gym and OpenHands.
- Controller: Google Cloud says Agent Sandbox Controller v1.0.0 limits how quickly the warm pool is refilled. This keeps etcd and the API server stable. Both are core Kubernetes components that track and manage the cluster.
- Customer: Google Cloud says the component already powers the agentic RL training infrastructure at Mistral AI, which it calls a frontier AI lab.
The benchmark figures Google Cloud published
Google Cloud says it tested on a cluster of 10 machines (nodes) running gVisor sandboxes, with GKE Image Streaming turned on. The tests used two workloads: a SWE-bench coding environment with 500 images, and an R2E collection with 4,578 images. Between 500 and 18,312 tasks ran at the same time. The R2E-Gym-Subset dataset page on Hugging Face lists 4,578 rows, which matches the size Google Cloud gives. That page says nothing about performance.
| Metric | Standard Kubernetes pods | GKE Agent Sandbox SDK | Improvement claimed by Google Cloud |
|---|---|---|---|
| Time to first command (average) | 44 to 85 seconds | 1.1 to 8.8 seconds | About 10x |
| Worst-case wait time | 7.5 minutes (450 seconds) | Under 10 seconds | About 45x |
| Pods created (4,578 images × 4 rollouts) | 18,312 | 5,869 | 3.1x less churn |
Google Cloud credits two changes for the speed-up. First, the software images are made ready ahead of time rather than while training waits. Second, the pace at which the cluster's control plane handles requests is regulated. The control plane is the part of Kubernetes that manages everything else.
A pod is the basic unit Kubernetes runs, and a rollout is one attempt by the agent at a task. Google Cloud says fewer pods are created because the SDK reuses the same pods between rollouts instead of deleting and rebuilding them. Google Cloud also says its benchmarking tools and load tests come with the agent-sandbox-rl example, so teams can repeat the comparison on their own clusters.
Why this matters
This is a change to the behind-the-scenes infrastructure that AI companies use. Many AI agents are now trained to write code and use tools, and that training means running attempts again and again in securely isolated environments. Google Cloud points out a bottleneck in this process: in a synchronous RL training step, training cannot continue until the slowest sandbox in the batch is ready. So slow-starting sandboxes leave expensive GPUs waiting.
If Google Cloud's claims hold, AI labs could do more agent training and testing with less wasted computing power. That could indirectly speed up work on products such as AI coding assistants. Google Cloud also quotes Mistral AI research engineer Jean-Malo Delignon as saying they can handle spikes of more than 30,000 sandboxes on a single cluster.
Trade-offs and limitations
Google Cloud acknowledges a trade-off: the approach spends cheaper CPU and background cluster time preparing environments in advance, so the GPUs are not left idle. It also says the SDK counts every failed sandbox and marks it for retry rather than quietly dropping it, because untracked failures can skew the rewards the model learns from. The update does not directly change any apps that everyday users have. It mainly affects research teams and startups that train AI agents at large scale.
Frequently asked questions
What is GKE Agent Sandbox, in plain terms?
It is a Google Cloud tool for running AI agents' code in isolated, secure environments on Kubernetes. It keeps a pool of environments started in advance, so they are ready immediately. According to Google Cloud, the new release is tuned for reinforcement learning and comes with a Python toolkit and integrations with popular training tools.
Where do the "10x" and "45x" figures come from?
"45x" refers to the worst case. Google Cloud says the longest wait for a sandbox fell from 7.5 minutes (450 seconds) to under 10 seconds. "10x" refers to the average time to run a first command, which Google Cloud says fell from 44–85 seconds to 1.1–8.8 seconds. Both come from Google Cloud's own tests on a 10-machine cluster.
Has anyone else confirmed these figures?
Not yet. The figures come only from Google Cloud's blog. Google Cloud says teams can rerun the comparison on their own clusters using the agent-sandbox-rl example.
Does this affect everyday users?
Not directly. It is cloud infrastructure for AI labs and startups that train AI agents. If it works as Google Cloud describes, it could help those teams train and test agents more efficiently, but there is no information on any specific effect for individual users.
Who is using it?
Google Cloud names Mistral AI, which it says uses the component for its agentic RL training infrastructure. The post also quotes one of Mistral AI's research engineers. No other customers are named.
Browse the latest news in this topic
Lifestyle
Cloudflare open-sources Streamline: a demo of using its cloud services to add graphics to live streams and burn subtitles into videos
On October 2, 2026, Cloudflare launched and open-sourced Streamline, a developer playground showing how developers can combine Stream, Workers, Containers and Durable Objects to build their own video processing pipelines, such as adding graphics to live streams in real time or adding subtitles to videos. This article explains what it is, how it works, its limitations, and what it means for viewers and developers.
Lifestyle
Google unveils Gemini 4 Argon: cyber defenders get it first, everyone else still has to wait
On September 30, 2026, Google announced Gemini 4 Argon, which it calls its new frontier (most advanced) AI model. For now it is available only to trusted cyber defenders through the Fairwind Program. Here is what Google says the model can do, what it will cost developers, how Google says it is managing the risks, and what it means for everyday users. All figures come from Google itself.
Lifestyle
NVIDIA: CoreWeave Begins Offering Vera Rubin NVL72, With Cognition as First Production Customer
According to the NVIDIA blog, AI cloud provider CoreWeave now offers NVIDIA's next-generation Vera Rubin NVL72 systems, plans to offer the Vera CPU and launched CoreWeave Forge. This matters mainly to companies building AI agents, and could eventually mean faster AI tools for everyday users.
Lifestyle
Google Cloud makes Spanner Omni generally available: its Spanner database can now run in companies' own data centers and on other clouds
Google Cloud says Spanner Omni, a version of its Spanner database that businesses run themselves, is now ready for real-world use in their own data centers, on other clouds or on a laptop. This matters to organizations that want Spanner outside Google Cloud, but they take on the running of it. Here are the features, licences and trade-offs Google Cloud describes.
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family
Sources
- Accelerate agentic RL with GKE Agent Sandbox | Google Cloud Blog · Checked:
- R2E-Gym/R2E-Gym-Subset · Datasets at Hugging Face · Checked:
- R2E-Gym (R2E-Gym) · Checked: