Lifestyle
NVIDIA: CoreWeave Begins Offering Vera Rubin NVL72, With Cognition as First Production Customer
According to the NVIDIA blog, AI cloud provider CoreWeave now offers NVIDIA's next-generation Vera Rubin NVL72 systems, plans to offer the Vera CPU and launched CoreWeave Forge. This matters mainly to companies building AI agents, and could eventually mean faster AI tools for everyday users.
About 6 min read

What happened
In a post dated September 30, 2026, the NVIDIA blog said CoreWeave made the announcement at its CoreWeave Fully Connected event in San Francisco, held the week of that post. The blog describes CoreWeave as a cloud purpose-built for AI. CoreWeave announced availability of NVIDIA Vera Rubin NVL72 systems paired with Spectrum-X 102.4T Ethernet networking. NVIDIA said this makes CoreWeave one of the first cloud providers to put the platform in customers' hands.
According to the NVIDIA blog, Cognition is the first customer running production workloads on Vera Rubin, meaning real work for real users rather than tests. Cognition is the applied AI lab behind the AI software engineer Devin. NVIDIA said CoreWeave received its first Vera Rubin NVL72 production racks earlier that month and stood up a production cluster for Cognition within days.
Read the full description
Sources are collected, independently checked, then reviewed by Jev.
NVIDIA also said CoreWeave will offer the NVIDIA Vera CPU and is launching CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing. AI agents are AI systems that carry out multistep tasks, such as writing and testing code. NVIDIA Vice President Ian Buck noted that CoreWeave's V100 GPUs are still running customer workloads nearly a decade after Volta launched.
Performance figures published by NVIDIA
Inference is the work of running a trained AI model to produce answers. Tokens are the small chunks of text a model reads and writes, and throughput is how many it can process. According to the NVIDIA blog, Cognition compared Vera Rubin NVL72 against a GB200 NVL72 baseline. It used a real-world software engineering workload made by sampling tasks from FrontierCode and having AI agents solve them. Early tests showed Vera Rubin NVL72 delivering up to 4.8x higher total token throughput on the SWE-2 inference workload. Silas Alberti, a member of Cognition's founding team, said agentic coding involves long contexts, high concurrency and large token volumes, and that the cost per token decides what they can ship.
On the CPU side, NVIDIA said CoreWeave's Vera deployment fits 128 CPUs and 11,264 cores in a single rack. At one core each, that is enough for more than 11,000 concurrent environments. These environments are sandboxes: isolated spaces where an agent can run code without touching other work. In testing, CoreWeave achieved more than 3x faster agent sandbox startup. It also saw a 1.7x performance gain across all passing Terminal-Bench tasks.
| Item | What the NVIDIA blog says | Status |
|---|---|---|
| Vera Rubin NVL72 | Up to 4.8x total token throughput on SWE-2 inference vs. GB200 NVL72 (Cognition early testing) | Announced as available on CoreWeave Cloud |
| NVIDIA Vera CPU | 128 CPUs and 11,264 cores per rack; sandbox startup more than 3x faster | Coming to CoreWeave Cloud |
| CoreWeave ARIA | Helps analyze experiments and propose experiments and code changes | Generally available |
| CoreWeave Sandboxes | Runs agents, RL and evaluations in isolated CPU or GPU environments | Generally available |
| CoreWeave Agent Lens | 20% better failure detection, fixes issues at half the cost (official claim) | New service |
| RL Rollouts | Loads new checkpoints while running, without redeployment | Private preview |
What is CoreWeave Forge
According to the NVIDIA blog, CoreWeave Forge brings together three things in one connected environment. They are Weights & Biases, post-training expertise from OpenPipe and the open-source marimo notebook project. Post-training means improving a model after its initial training. The goal is to let how models and agents behave in real use feed into the next round of training, while staying open across models, frameworks and clouds. ARIA and Sandboxes are generally available, and Agent Lens is a new service.
Reinforcement learning (RL) is a training method in which a model improves through trial and feedback. NVIDIA said serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup. The open-source inference framework NVIDIA Dynamo powers CoreWeave's managed inference service and RL Rollouts, which is in private preview. RL Rollouts loads new checkpoints, which are saved versions of a model, into a live deployment. According to NVIDIA, Canva, Capital One and MasterClass are among the first companies building on Forge.
What it means for general readers
The key point is that a new generation of AI hardware is starting to enter cloud production environments. Agentic services such as AI coding assistants need large amounts of computing power. If the throughput gains NVIDIA describes hold up in real use, end users may notice faster responses and smoother multistep reasoning. NVIDIA also gave a healthcare example. According to NVIDIA, Ennoble Care serves about 50,000 high-need Medicare patients across 15 states. It will use RTX PRO 6000 GPUs on CoreWeave to scale AI agents for clinical documentation, decision support and back-office automation.
- General users don't need to do anything; this is a cloud and enterprise-level infrastructure update.
- For developers, the general availability of isolated sandboxes and post-training tools may lower the barrier to building AI agents, but rely on your own testing.
- The performance figures come from the vendor and its partners; it's worth watching whether independent benchmarks follow.
Frequently asked questions
What is Vera Rubin NVL72?
According to the NVIDIA blog, it is NVIDIA's next-generation AI infrastructure system. CoreWeave announced it in its cloud, paired with Spectrum-X 102.4T Ethernet networking.
Is the 4.8x performance gain credible?
It is a result NVIDIA cited from Cognition's early testing. It is the best case ("up to") on one workload, SWE-2 inference, measured against a GB200 NVL72 system. It has not been independently verified.
Who is the first user?
According to the NVIDIA blog, Cognition, the developer of Devin, is the first customer running production workloads on Vera Rubin.
How is the Vera CPU different from a regular CPU?
NVIDIA calls Vera the first CPU built for AI agents, focused on running large numbers of isolated agent environments at once. CoreWeave's deployment has 11,264 cores per rack.
Will this affect the AI services I use every day?
Possibly, indirectly. If infrastructure performance improves as NVIDIA describes, AI apps using these cloud services may respond faster, but actual results depend on each provider's deployment.
Browse the latest news in this topic
Lifestyle
Cloudflare open-sources Streamline: a demo of using its cloud services to add graphics to live streams and burn subtitles into videos
On October 2, 2026, Cloudflare launched and open-sourced Streamline, a developer playground showing how developers can combine Stream, Workers, Containers and Durable Objects to build their own video processing pipelines, such as adding graphics to live streams in real time or adding subtitles to videos. This article explains what it is, how it works, its limitations, and what it means for viewers and developers.
Lifestyle
Google unveils Gemini 4 Argon: cyber defenders get it first, everyone else still has to wait
On September 30, 2026, Google announced Gemini 4 Argon, which it calls its new frontier (most advanced) AI model. For now it is available only to trusted cyber defenders through the Fairwind Program. Here is what Google says the model can do, what it will cost developers, how Google says it is managing the risks, and what it means for everyday users. All figures come from Google itself.
Lifestyle
Google Cloud makes Spanner Omni generally available: its Spanner database can now run in companies' own data centers and on other clouds
Google Cloud says Spanner Omni, a version of its Spanner database that businesses run themselves, is now ready for real-world use in their own data centers, on other clouds or on a laptop. This matters to organizations that want Spanner outside Google Cloud, but they take on the running of it. Here are the features, licences and trade-offs Google Cloud describes.
Lifestyle
Google Cloud Previews a Remote MCP Server That Lets AI Agents Run gcloud and bq Commands for You
Google Cloud has opened a public preview of its Google Cloud CLI remote MCP server. It lets AI agents run Google Cloud's gcloud and bq command-line tools in isolated cloud sandboxes, with nothing to install locally. This article explains how it works, the safeguards Google describes, the pricing, and what it means for businesses, developers and everyday readers.
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family