Lifestyle

Google Cloud says its new GKE Agent Sandbox starts AI training sandboxes 10 to 45 times faster

Google Cloud has released a version of its GKE Agent Sandbox built for training AI agents, plus a Python toolkit and ready-made integrations. It says the sandboxes start 10 to 45 times faster, so costly GPUs spend less time sitting idle. The main audience is AI labs and startups. The speed figures come only from Google Cloud's own tests.

About 7 min read

Google Cloud says its new GKE Agent Sandbox starts AI training sandboxes 10 to 45 times faster
Image: Mokaair (Original editorial artwork)

What Google Cloud announced

Google Cloud made the announcement in a blog post dated September 30, 2026. It covers three things: a version of GKE Agent Sandbox optimized for reinforcement learning, the Agent Sandbox RL orchestration SDK (a software toolkit for managing many sandboxes at once), and native integrations with common RL "gyms" and tools, which are ready-made training environments. Google Cloud says all three are now generally available. The post was written by product manager Tinsley Shi and staff software engineer and technical lead Tomer Glottmann.

Google Cloud describes the problem this way. When an AI agent is trained with RL, a model running on GPUs (costly AI chips) produces actions such as pieces of code. Those actions are then run in a sandbox, an isolated computer environment on ordinary CPUs, so the result can be scored as feedback. At large scale, the GPUs sit idle while they wait. They wait for sandboxes to cold start (boot from scratch), for many multi-gigabyte container images (packaged copies of the software each task needs) to download, and for scheduling queues to clear.

Google Cloud says its new GKE Agent Sandbox starts AI training sandboxes 10 to 45 times faster
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

GKE is Google Cloud's service for running Kubernetes, a widely used open system for running software across many machines. Google Cloud describes GKE Agent Sandbox as an open Kubernetes building block for running agents securely. It includes a built-in SandboxWarmPool, a "warm pool" of environments started in advance so they are ready to use. It also supports snapshots, pause and resume.

  • Orchestration SDK: Google Cloud says it offers an asynchronous Python API, so researchers do not have to write Kubernetes configuration files (YAML).
  • Tool integrations: Google Cloud lists Gymnasium, NVIDIA NeMo Gym and OpenHands.
  • Controller: Google Cloud says Agent Sandbox Controller v1.0.0 limits how quickly the warm pool is refilled. This keeps etcd and the API server stable. Both are core Kubernetes components that track and manage the cluster.
  • Customer: Google Cloud says the component already powers the agentic RL training infrastructure at Mistral AI, which it calls a frontier AI lab.

The benchmark figures Google Cloud published

Google Cloud says it tested on a cluster of 10 machines (nodes) running gVisor sandboxes, with GKE Image Streaming turned on. The tests used two workloads: a SWE-bench coding environment with 500 images, and an R2E collection with 4,578 images. Between 500 and 18,312 tasks ran at the same time. The R2E-Gym-Subset dataset page on Hugging Face lists 4,578 rows, which matches the size Google Cloud gives. That page says nothing about performance.

Source: Google Cloud blog; results from Google Cloud's own tests
MetricStandard Kubernetes podsGKE Agent Sandbox SDKImprovement claimed by Google Cloud
Time to first command (average)44 to 85 seconds1.1 to 8.8 secondsAbout 10x
Worst-case wait time7.5 minutes (450 seconds)Under 10 secondsAbout 45x
Pods created (4,578 images × 4 rollouts)18,3125,8693.1x less churn

Google Cloud credits two changes for the speed-up. First, the software images are made ready ahead of time rather than while training waits. Second, the pace at which the cluster's control plane handles requests is regulated. The control plane is the part of Kubernetes that manages everything else.

A pod is the basic unit Kubernetes runs, and a rollout is one attempt by the agent at a task. Google Cloud says fewer pods are created because the SDK reuses the same pods between rollouts instead of deleting and rebuilding them. Google Cloud also says its benchmarking tools and load tests come with the agent-sandbox-rl example, so teams can repeat the comparison on their own clusters.

Why this matters

This is a change to the behind-the-scenes infrastructure that AI companies use. Many AI agents are now trained to write code and use tools, and that training means running attempts again and again in securely isolated environments. Google Cloud points out a bottleneck in this process: in a synchronous RL training step, training cannot continue until the slowest sandbox in the batch is ready. So slow-starting sandboxes leave expensive GPUs waiting.

If Google Cloud's claims hold, AI labs could do more agent training and testing with less wasted computing power. That could indirectly speed up work on products such as AI coding assistants. Google Cloud also quotes Mistral AI research engineer Jean-Malo Delignon as saying they can handle spikes of more than 30,000 sandboxes on a single cluster.

Trade-offs and limitations

Google Cloud acknowledges a trade-off: the approach spends cheaper CPU and background cluster time preparing environments in advance, so the GPUs are not left idle. It also says the SDK counts every failed sandbox and marks it for retry rather than quietly dropping it, because untracked failures can skew the rewards the model learns from. The update does not directly change any apps that everyday users have. It mainly affects research teams and startups that train AI agents at large scale.

Frequently asked questions

What is GKE Agent Sandbox, in plain terms?

It is a Google Cloud tool for running AI agents' code in isolated, secure environments on Kubernetes. It keeps a pool of environments started in advance, so they are ready immediately. According to Google Cloud, the new release is tuned for reinforcement learning and comes with a Python toolkit and integrations with popular training tools.

Where do the "10x" and "45x" figures come from?

"45x" refers to the worst case. Google Cloud says the longest wait for a sandbox fell from 7.5 minutes (450 seconds) to under 10 seconds. "10x" refers to the average time to run a first command, which Google Cloud says fell from 44–85 seconds to 1.1–8.8 seconds. Both come from Google Cloud's own tests on a 10-machine cluster.

Has anyone else confirmed these figures?

Not yet. The figures come only from Google Cloud's blog. Google Cloud says teams can rerun the comparison on their own clusters using the agent-sandbox-rl example.

Does this affect everyday users?

Not directly. It is cloud infrastructure for AI labs and startups that train AI agents. If it works as Google Cloud describes, it could help those teams train and test agents more efficiently, but there is no information on any specific effect for individual users.

Who is using it?

Google Cloud names Mistral AI, which it says uses the component for its agentic RL training infrastructure. The post also quotes one of Mistral AI's research engineers. No other customers are named.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle