Lifestyle

NVIDIA: CoreWeave Begins Offering Vera Rubin NVL72, With Cognition as First Production Customer

According to the NVIDIA blog, AI cloud provider CoreWeave now offers NVIDIA's next-generation Vera Rubin NVL72 systems, plans to offer the Vera CPU and launched CoreWeave Forge. This matters mainly to companies building AI agents, and could eventually mean faster AI tools for everyday users.

About 6 min read

NVIDIA: CoreWeave Begins Offering Vera Rubin NVL72, With Cognition as First Production Customer
Image: Mokaair (Original editorial artwork)

What happened

In a post dated September 30, 2026, the NVIDIA blog said CoreWeave made the announcement at its CoreWeave Fully Connected event in San Francisco, held the week of that post. The blog describes CoreWeave as a cloud purpose-built for AI. CoreWeave announced availability of NVIDIA Vera Rubin NVL72 systems paired with Spectrum-X 102.4T Ethernet networking. NVIDIA said this makes CoreWeave one of the first cloud providers to put the platform in customers' hands.

According to the NVIDIA blog, Cognition is the first customer running production workloads on Vera Rubin, meaning real work for real users rather than tests. Cognition is the applied AI lab behind the AI software engineer Devin. NVIDIA said CoreWeave received its first Vera Rubin NVL72 production racks earlier that month and stood up a production cluster for Cognition within days.

NVIDIA: CoreWeave Begins Offering Vera Rubin NVL72, With Cognition as First Production Customer
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

NVIDIA also said CoreWeave will offer the NVIDIA Vera CPU and is launching CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing. AI agents are AI systems that carry out multistep tasks, such as writing and testing code. NVIDIA Vice President Ian Buck noted that CoreWeave's V100 GPUs are still running customer workloads nearly a decade after Volta launched.

Performance figures published by NVIDIA

Inference is the work of running a trained AI model to produce answers. Tokens are the small chunks of text a model reads and writes, and throughput is how many it can process. According to the NVIDIA blog, Cognition compared Vera Rubin NVL72 against a GB200 NVL72 baseline. It used a real-world software engineering workload made by sampling tasks from FrontierCode and having AI agents solve them. Early tests showed Vera Rubin NVL72 delivering up to 4.8x higher total token throughput on the SWE-2 inference workload. Silas Alberti, a member of Cognition's founding team, said agentic coding involves long contexts, high concurrency and large token volumes, and that the cost per token decides what they can ship.

On the CPU side, NVIDIA said CoreWeave's Vera deployment fits 128 CPUs and 11,264 cores in a single rack. At one core each, that is enough for more than 11,000 concurrent environments. These environments are sandboxes: isolated spaces where an agent can run code without touching other work. In testing, CoreWeave achieved more than 3x faster agent sandbox startup. It also saw a 1.7x performance gain across all passing Terminal-Bench tasks.

CoreWeave's new services and NVIDIA's published figures (source: NVIDIA blog)
ItemWhat the NVIDIA blog saysStatus
Vera Rubin NVL72Up to 4.8x total token throughput on SWE-2 inference vs. GB200 NVL72 (Cognition early testing)Announced as available on CoreWeave Cloud
NVIDIA Vera CPU128 CPUs and 11,264 cores per rack; sandbox startup more than 3x fasterComing to CoreWeave Cloud
CoreWeave ARIAHelps analyze experiments and propose experiments and code changesGenerally available
CoreWeave SandboxesRuns agents, RL and evaluations in isolated CPU or GPU environmentsGenerally available
CoreWeave Agent Lens20% better failure detection, fixes issues at half the cost (official claim)New service
RL RolloutsLoads new checkpoints while running, without redeploymentPrivate preview

What is CoreWeave Forge

According to the NVIDIA blog, CoreWeave Forge brings together three things in one connected environment. They are Weights & Biases, post-training expertise from OpenPipe and the open-source marimo notebook project. Post-training means improving a model after its initial training. The goal is to let how models and agents behave in real use feed into the next round of training, while staying open across models, frameworks and clouds. ARIA and Sandboxes are generally available, and Agent Lens is a new service.

Reinforcement learning (RL) is a training method in which a model improves through trial and feedback. NVIDIA said serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup. The open-source inference framework NVIDIA Dynamo powers CoreWeave's managed inference service and RL Rollouts, which is in private preview. RL Rollouts loads new checkpoints, which are saved versions of a model, into a live deployment. According to NVIDIA, Canva, Capital One and MasterClass are among the first companies building on Forge.

What it means for general readers

The key point is that a new generation of AI hardware is starting to enter cloud production environments. Agentic services such as AI coding assistants need large amounts of computing power. If the throughput gains NVIDIA describes hold up in real use, end users may notice faster responses and smoother multistep reasoning. NVIDIA also gave a healthcare example. According to NVIDIA, Ennoble Care serves about 50,000 high-need Medicare patients across 15 states. It will use RTX PRO 6000 GPUs on CoreWeave to scale AI agents for clinical documentation, decision support and back-office automation.

  • General users don't need to do anything; this is a cloud and enterprise-level infrastructure update.
  • For developers, the general availability of isolated sandboxes and post-training tools may lower the barrier to building AI agents, but rely on your own testing.
  • The performance figures come from the vendor and its partners; it's worth watching whether independent benchmarks follow.

Frequently asked questions

What is Vera Rubin NVL72?

According to the NVIDIA blog, it is NVIDIA's next-generation AI infrastructure system. CoreWeave announced it in its cloud, paired with Spectrum-X 102.4T Ethernet networking.

Is the 4.8x performance gain credible?

It is a result NVIDIA cited from Cognition's early testing. It is the best case ("up to") on one workload, SWE-2 inference, measured against a GB200 NVL72 system. It has not been independently verified.

Who is the first user?

According to the NVIDIA blog, Cognition, the developer of Devin, is the first customer running production workloads on Vera Rubin.

How is the Vera CPU different from a regular CPU?

NVIDIA calls Vera the first CPU built for AI agents, focused on running large numbers of isolated agent environments at once. CoreWeave's deployment has 11,264 cores per rack.

Will this affect the AI services I use every day?

Possibly, indirectly. If infrastructure performance improves as NVIDIA describes, AI apps using these cloud services may respond faster, but actual results depend on each provider's deployment.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle