Lifestyle

NVIDIA's AI Infra Summit Power Efficiency Talk: What's New With Vera Rubin and DSX

On September 15, 2026, at the AI Infra Summit in Santa Clara, California, NVIDIA presented the Vera Rubin platform and DSX data-center power-management software, saying the metric for AI infrastructure is shifting from peak performance to tokens produced per megawatt. Based on NVIDIA's blog posts from that day and its January and August press releases, this article checks what was actually new, which numbers are measured, which NVIDIA labels as projections, and what remains unstated.

About 13 min read

Original illustration: a chip, a rack, and a data-center building linked by arrows, then connected by a dashed line to a transmission tower, showing four layers: chip, rack, data center, and grid.
Image: Mokaair (© Mokaair)

On September 15, 2026, according to a roundup post NVIDIA's official blog published that same day, NVIDIA vice president Ian Buck spoke on AI factory efficiency at the AI Infra Summit, held at the Santa Clara Convention Center in California, making it the centerpiece of his infrastructure keynote.

This article was fact-checked on September 17, 2026, against the full text of the two official NVIDIA blog posts published that day, plus two official press releases from January 5 and August 24 used as a baseline for comparison. This site has not run any tests of its own, and offers no purchasing or upgrade advice; every number below is labeled as either something NVIDIA measured itself, something NVIDIA projects, or a result NVIDIA attributes to a third-party testing organization.

What Happened on September 15: From Peak Performance to Tokens per Megawatt

NVIDIA states in this roundup post that the metric for AI infrastructure is quickly shifting from peak performance to validated agentic tokens per megawatt, and that AI factories must now be codesigned from silicon to grid. That statement is what separates September 15 from the January CES announcement: January was about chip specifications and the cost per token, while September is about how much output an entire data center can squeeze from each megawatt of power.

Neither of the two September 15 announcements introduced a new chip; instead, NVIDIA listed ecosystem partnerships and partner test results. NVIDIA says Amazon's Annapurna Labs is working with it on the NVHBM custom high-bandwidth memory technology; d-Matrix is integrating NVLink Fusion to combine NVIDIA Vera CPUs with its own Raptor XPUs for ultra-low-latency inference at scale; and Pinterest is using the Blackwell platform and Dynamo inference software to bring conversational AI to visual discovery. NVIDIA gives each of these three items only a single sentence, with no specifications or timeline; the same post also includes Vera CPU test results from several startups and a section on NVLink 6 reliability.

Compared With January: The Chips Haven't Changed, the Metric Has

On January 5, NVIDIA unveiled the Rubin platform at CES, describing it as six new chips, including the Vera CPU and Rubin GPU. The headline numbers were about performance: up to a 10x reduction in per-token inference cost and a 4x reduction in the number of GPUs needed to train mixture-of-experts models, both compared with the Blackwell platform; the press release used future tense, saying partner products were expected in the second half of 2026. At the Hot Chips conference on August 24, NVIDIA announced that Groq 3 LPX had reached full production, describing the platform as spanning seven chips and five purpose-built racks; the only trademark note states that the Groq and LPU names are used under license from Groq, Inc., and the official documents checked for this article make no mention of NVIDIA acquiring or merging with Groq.

By September 15, neither official blog post introduced any new chip name; the January press release never used megawatts as a unit of measurement anywhere in its text, while the two September 15 posts repeatedly describe results in terms of tokens produced per megawatt. In other words, what's new on September 15 isn't hardware — it's that the DSX power-management software has reached the stage of its first round of partner testing and commercial demonstrations.

Compiled from NVIDIA announcements published January 5, August 24, September 15, and September 16, 2026. Fact-checked September 17, 2026.
DateWhat the Announcement CoveredThe Numbers and Conditions NVIDIA Gave
2026-01-05CES: Rubin six-chip platform unveiledNVIDIA said up to a 10x cut in per-token inference cost and 4x fewer GPUs for training, vs. Blackwell
2026-08-24Hot Chips: Groq 3 LPX reaches full productionNVIDIA said the platform spans seven chips and five racks in total; Groq and LPU are licensed from Groq, Inc.
2026-09-15AI Infra Summit: DSX power software and partner testsNVIDIA said Lambda measured 24% more throughput; the 40% figure for Vera Rubin NVL72 is NVIDIA's own projection
2026-09-16Same post updated with MLPerf Inference v6.1 preview resultsNVIDIA said Vera Rubin NVL72 throughput is up to 3.7x that of GB300 NVL72

Why Power Is the Bottleneck: What DSX MaxLPS and DSX Flex Do

NVIDIA quotes founder and CEO Jensen Huang to explain why power is the bottleneck: “A one-gigawatt factory will never become a two-gigawatt factory,” Huang has said. NVIDIA writes in the same post that power is the constraint in the AI factory economy. DSX software was not launched on September 15 — NVIDIA states that DSX was introduced at GTC Taipei, an NVIDIA event held in Taiwan, in May, without giving a year; neither of the two September 15 announcements says anything further about deployments, partners, or supply chains in Taiwan, and this article does not extend the record on NVIDIA's behalf.

DSX MaxLPS handles power dispatch within a data center. NVIDIA says that at the factory level, MaxLPS dynamically shifts power across racks as demand rises and falls, delivering up to 40% more GPUs within the same site-power envelope, and up to 35% higher token throughput without requiring new power lines; at the rack level, separate power-smoothing software and expanded energy buffering absorb short spikes, though NVIDIA gives no percentage figure at that level. The one partner-measured result in this set of power figures is Lambda's: running 19 nodes under an 85% power policy within the power budget of 16 nodes at full power, NVIDIA says the cluster measured 24% more throughput and 23% better performance per watt — but that test ran on Blackwell-generation HGX B200 servers, not Vera Rubin, and Lambda's president of cloud services calls it a proof of concept. When that same 40% figure is attached to the Vera Rubin NVL72, the companion post states that it is based on NVIDIA's own projections and limited to suitable deployment environments; that post's figures table further notes it requires pairing MaxLPS with data center power planning, and is not a measured result.

DSX Flex handles communication between a data center and the power grid: it receives load-shedding requests, demand-response events, and pricing signals from the grid, and reacts automatically within a predefined workload hierarchy, keeping the most critical jobs running while everything else pauses and later resumes. NVIDIA's example is its own Eos data center, which took part in a flexible load interconnect program with Silicon Valley Power, paired with Emerald AI's Conductor: the text says power automatically dropped from 4 megawatts to 3 megawatts one August evening, and a photo caption on the same page states the same two figures; but a separate figures table in the same post describes that same demonstration as a 40% reduction in power demand within one minute, and NVIDIA does not explain the relationship between the two. NVIDIA says Silicon Valley Power has since sent more than 200 demand signals, every one of them successful, but this demonstration is not yet a formal DSX Flex installation — only proof that the concept works.

Four-panel diagram: power and efficiency dispatch at four layers — chip, rack, data center, and grid
Compiled from two NVIDIA blog posts published September 15, 2026. Fact-checked September 17, 2026. · Image: Mokaair (© Mokaair)

How to Read NVIDIA's Numbers: What's Measured, What's Projected

With Groq 3 LPX attached, NVIDIA's headline number is that for models with more than two trillion parameters at long context, the combined platform delivers up to 35 times the token throughput per megawatt of the GB200 NVL72; a more specific figure is that on a 100,000-token-context Qwen 3.8 27B workload, Groq 3 LPX measured 2,529 output tokens per second per user — both figures carry their own comparison baseline, and quoting either without its conditions would distort it. NVIDIA also cites a result from third-party firm SemiAnalysis, measured with its AgentX methodology: on the DeepSeek V4 Pro model, the Vera Rubin NVL72 delivers up to 30 times the throughput per megawatt of the GB300 NVL72 — this is NVIDIA relaying another organization's test, and this site has not obtained the underlying data.

On availability, the two September 15 blog posts give only one sentence: Vera Rubin is in full production and scaling across the ecosystem, and performance will keep improving with ongoing software optimization, with no shipment volume, customer list, or list of cloud regions. This particular post is a living document that keeps being updated: on September 16, the day after the event, it added a section of MLPerf Inference v6.1 results, which NVIDIA describes as the Vera Rubin NVL72's first MLPerf preview submission, delivering up to 3.7 times the throughput of the GB300 NVL72. That content belongs to September 16, not to what happened on September 15 itself.

What This Means for Cloud Service Users, and What NVIDIA Didn't Say

Every number above is a data-center-level metric — tokens produced per megawatt, performance per watt, how many GPUs fit within a given power allotment — and none of them is a promise about end-user cloud service speed, pricing, or delivery timing. Whether hardware efficiency ever shows up in a lower monthly subscription bill is a gap between inference cost and end-user pricing this site's earlier CES coverage already broke down, and it will not be repeated here; neither of the two September 15 announcements mentions a price or licensing fee for Vera Rubin, Groq 3 LPX, or DSX at any point, and the word “price” only appears when describing grid pricing signals.

As checked on September 17, 2026, neither post contains the words “Taiwan” or “TSMC”; there is no description of any Taiwan deployment, Taiwan partner, or supply chain. The only Taiwan connection is the one noted in the previous section — that DSX was introduced at GTC Taipei. Readers who want to confirm the current state for themselves can go back to the two blog posts directly; the summit post's last-updated timestamp moved forward again during the course of fact-checking. If readers later see an updated version, they should treat whatever the official page says at that time as authoritative.

Frequently asked questions

Did NVIDIA announce a new chip at this AI Infra Summit?

No. The Vera Rubin platform was unveiled on January 5, 2026, and Groq 3 LPX was announced as reaching full production on August 24. On September 15, NVIDIA's official blog posts covered power-management software, the first round of partner test results, and test results several startups ran on the Vera CPU; neither of the two announcements checked for this article contains a new chip name.

Does DSX MaxLPS's 40% more GPUs for a Vera Rubin data center come from a measurement?

No. That 40% figure refers to fitting 40% more GPUs within the same power budget, and in the official post “From Megawatts to Tokens,” NVIDIA labels it as based on NVIDIA's own projections and limited to suitable deployment environments; that post's figures table further states it is the result of pairing MaxLPS with data center power planning. The same claim, made in the summit roundup post published the same day, is not labeled as a projection there. The one measured figure between the two posts is Lambda's test, which found 24% more throughput and 23% better performance per watt, but that ran on Blackwell-generation HGX B200 servers, not Vera Rubin, and Lambda's president of cloud services calls it a proof of concept.

In the Silicon Valley Power load-shedding demonstration, how much did power actually drop?

NVIDIA describes it two different ways in the same post: the body text and a photo caption on the same page both say power automatically dropped from 4 megawatts to 3 megawatts, while a separate figures table describes the same demonstration as a 40% reduction in power demand within one minute. NVIDIA does not explain the relationship between the two figures itself, and this article lists both rather than choosing one as the single answer.

Does this have anything to do with Taiwan?

The two September 15 announcements checked for this article do not mention Taiwan or TSMC anywhere. The only connection is that the DSX power-management software itself was introduced at GTC Taipei, an NVIDIA event held in Taiwan, in May; the announcements give no further detail about any Taiwan deployment, partner, or supply chain, and this article does not fill that in on NVIDIA's behalf.

Can you buy Vera Rubin now, and how much does it cost?

NVIDIA's own statement is that it is in full production and scaling across the ecosystem, but it has not disclosed shipment volumes, a customer list, or which cloud regions carry it; the January press release said partner products were expected in the second half of 2026, and the September 15 materials did not update that timeline. Neither of the two September 15 announcements checked for this article states a price or licensing fee for Vera Rubin, Groq 3 LPX, or DSX, and this article offers no advice on whether to adopt any of them.

What does the 30x figure from SemiAnalysis mean?

That's a result NVIDIA attributes to third-party testing firm SemiAnalysis, measured using its AgentX methodology: on the DeepSeek V4 Pro model, the Vera Rubin NVL72 delivers up to 30 times the throughput per megawatt of the GB300 NVL72. This is NVIDIA citing another organization's test; this site has not obtained SemiAnalysis's underlying data, and can only report how NVIDIA cites it.

Latest travel guides

Sources

Lifestyle