Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

NVIDIA says DSX boosted Lambda’s Blackwell throughput 24% within a fixed power budget

NVIDIA says Lambda used DSX MaxLPS to increase token throughput within a fixed power budget, while an Emerald AI system demonstrated automated load reduction in response to utility signals.

D
Sep 15, 2026 · 3 min read

NVIDIA said its DSX power-management software enabled Lambda to run more Blackwell systems within a fixed facility power budget, increasing token throughput by 24% in a test disclosed Sept. 15. The company also detailed a separate Emerald AI deployment that it says cut an AI factory’s load by one megawatt in under a minute in response to a utility signal.

The two results address different power constraints in AI infrastructure. MaxLPS adjusts GPU power allocations within a fixed facility limit, while DSX Flex is designed to coordinate workload priorities with signals from the electric grid. Both performance results come from NVIDIA and its partners rather than independently controlled tests.

In NVIDIA’s account of the Lambda test, 19 NVIDIA HGX B200 nodes operated under an 85% power policy within the same facility budget as a 16-node full-power baseline. NVIDIA and Lambda said pure-inference throughput rose from about 4.04 million to 5.00 million tokens per second, a 24% increase, while performance per watt improved 23%.

The Lambda case study says the proof of concept covered five racks and 19 HGX B200 nodes, using MLPerf inference and training workloads to produce consistent peak-level power draw. In a mixed-workload test, NVIDIA said an 80% policy across 10 inference nodes and 10 training nodes increased training-cluster throughput 17% and inference-cluster throughput 20% against the case study’s baseline. NVIDIA characterized the work as the first validation of DSX MaxLPS on Blackwell servers; the reviewed evidence does not independently establish that priority claim.

MaxLPS does not add electrical capacity. According to NVIDIA’s Dynamic Power Software documentation, it pools a GPU power budget derived from policy and uses live telemetry and workload-allocation data to adjust per-GPU limits while staying inside topology, hardware and policy constraints. The documentation says systems such as Slurm or Kubernetes still decide where jobs run, while NVIDIA’s software controls the GPU power available to them. That software layer sits alongside separate work on power delivery for denser AI racks.

NVIDIA also says the approach can support up to 40% more Vera Rubin NVL72 GPUs and up to 35% higher token throughput within the same site-power envelope. NVIDIA’s summit post gives those as ‘up to’ platform figures, while its documentation describes 40% more compute as a planning rule of thumb that operators should validate against their own workloads and power requirements. The reviewed pages do not provide a deployment test methodology for the Vera Rubin figures.

The grid-facing layer, DSX Flex, is designed to receive load-shedding, demand-response and pricing signals, then apply a predefined workload hierarchy. NVIDIA says that setup can pause or throttle lower-priority jobs while critical work continues.

The City of Santa Clara and Silicon Valley Power announced the underlying commercial, multi-megawatt pilot with Emerald AI in April. The announcement said Emerald AI’s Conductor software would respond to utility signals at a data center running NVIDIA workloads while protecting workload performance. It did not report the later operational results.

For an August event, NVIDIA said Conductor automatically reduced the load at its Eos AI factory from four megawatts to three in less than a minute, with low-priority jobs yielding while high-priority inference continued. NVIDIA also said the system later responded successfully to more than 200 Silicon Valley Power demand signals. No raw power trace, latency distribution or utility report independently confirming those figures was available in the research record.

A June research preprint provides separate evidence that grid-responsive computing can work at a smaller scale: it reports rapid load reduction and sustained curtailment in a 130-kilowatt GPU cluster while maintaining service levels for priority jobs. The paper does not verify the Lambda or Santa Clara results.

The partners describe the Santa Clara deployment differently. Emerald AI called it the first commercial DSX Flex deployment, but NVIDIA’s September account said it was not a DSX Flex installation and instead demonstrated behavior the companies plan to integrate. NVIDIA identified a planned facility in Manassas as the first dedicated commercial DSX Flex deployment.

More news