Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech.Before a packed audience — with more than 8,000 attendees this year, up from 3,500 last year — Buck discussed new collaborations across NVIDIA platforms and more.Amazon’s Annapurna Labs is working with NVIDIA on the NVHBM custom high-bandwidth memory technology.d-Matrix is integrating with NVLink Fusion to combine NVIDIA Vera CPUs with d-Matrix Raptor XPUs to deliver ultra low-latency inference at scale.NVIDIA and partners unveiled new results as well: Emerald AI and NVIDIA demonstrated a commercial AI factory flexible-load program, working with Silicon Valley Power.Lambda improved performance per watt by 23% with NVIDIA DSX MaxLPS.Pinterest is using the NVIDIA Blackwell platform and NVIDIA Dynamo inference software to bring conversational AI to visual discovery. The news comes as agentic AI is driving a new class of workloads that demand more performance, efficiency and scale from AI infrastructure. NVIDIA addresses that challenge with a full-stack AI factory platform spanning Vera Rubin systems, Dynamo inference software, NeMo libraries and NVIDIA networking — including NVIDIA NVLink for scale-up computing, Spectrum-X Ethernet and ConnectX SuperNICs for connecting thousands of nodes, BlueField-powered context-memory storage and BlueField DPUs for infrastructure security. The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt. AI factories must now be codesigned from silicon to grid. NVIDIA DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink helps unite large-scale accelerated computing into a single high-performance system. The result is AI infrastructure designed to generate more tokens, improve efficiency and help customers get more value from every megawatt of power. Tuesday, Sept. 15, 9:55 a.m. PT AI Factory Flexible-Load Program Unlocks More Grid PowerSilicon Valley Power operates a flexible-load interconnection program that enables AI factories to support grid flexibility. Through this program, Emerald AI worked with NVIDIA to demonstrate automated load reduction at Silicon Valley Power. The system successfully responded to hundreds of demand signals from Silicon Valley Power while protecting AI workload performance. NVIDIA partner Emerald AI is planning to use NVIDIA DSX Flex for its Conductor grid-responsive power management software that dynamically adjusts an AI factory’s energy consumption based on real-time electricity grid signals and hybrid energy sources. Using DSX Flex to autonomously load balance, Emerald AI can demonstrate its ability to automatically throttle back power on low-priority AI jobs and then return to normal. DSX Flex can receive grid signals — load-shedding requests, demand-response events and pricing signals — and automatically act within a predefined workload hierarchy. The most critical jobs keep going, while everything else pauses temporarily and then resumes. In this way, facilities can participate in energy efficiency for the grid. This capability enables AI factories to operate as flexible grid resources, so they can reduce demand when the grid needs relief, protect priority AI workloads and prove that controllable AI load can help unlock more grid capacity for growth. Learn more about the Emerald AI partnership with Silicon Valley Power. Tuesday, Sept. 15, 9:55 a.m. PT Lambda Maximizes Performance per Watt With NVIDIA DSX MaxLPSAI cloud provider Lambda released results at the AI Infra Summit, providing the first validation of NVIDIA DSX MaxLPS on NVIDIA Blackwell servers. DSX MaxLPS continuously monitors power consumption across GPUs and racks, dynamically shifting available power where it’s needed most and reclaiming capacity that static provisioning leaves unused. Because training and inference workloads have different power profiles, MaxLPS optimizes power allocation across mixed-workload AI factories. The results were significant: Lambda ran 19 nodes within the same power budget typically allocated to 16 full-power nodes, increasing cluster-wide token throughput by 24% — from about 4 million to 5 million tokens per second — while improving performance per watt by 23%. For next-generation NVIDIA Vera Rubin NVL72 AI factories, DSX MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget in the right deployment environments.This means AI factory operators can increase AI capacity and token throughput within the same power envelope, helping maximize the productivity and economic value of every megawatt.Learn more about Lambda’s results. Tuesday, Sept. 15, 9:55 a.m. PT Vera Rubin and Groq 3 LPX Turn More Power Into TokensIn AI factories, power is the constraint. Every watt counts, so squeezing more tokens from every megawatt is what matters most for AI infrastructure.NVIDIA does this with a full-stack AI factory built on the Vera Rubin NVL72 — systems, networking, software and power management working as one. At the factory level, NVIDIA DSX MaxLPS dynamically shifts power across racks as demand rises and falls. The payoff is significant: Up to 40% more GPUs within the same site-power envelope Up to 35% higher token throughput — no new power lines requiredInside each rack, Intelligent Power Smoothing software and expanded energy buffering absorb short spikes, letting systems run closer to sustained demand. Stranded headroom becomes productive compute.For agentic AI — where agents chain reasoning steps and tool calls — latency and context length compounds quickly. That’s where NVIDIA Groq 3 LPX comes in, adding deterministic ultralow-latency inference to Vera Rubin and complements DSX MaxLPS. The combined platform delivers up to 35X higher token throughput per megawatt than GB200 NVL72 for 2-trillion-plus-parameter models at long context.The numbers tell the story. On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX hit 2,529 output tokens per second per user. That headroom lets agents run more reasoning steps and tool calls within the same response budget, even as workloads scale.Learn more about how NVIDIA Vera Rubin NVL72 and Groq 3 LPX can produce more tokens with less power. Tuesday, Sept. 15, 9:55 a.m. PT Vera Rubin NVL72 Delivers 30x Better AI Factory Throughput on SemiAnalysis AgentXSemiAnalysis AgentX measures inference on recorded real-world agentic coding sessions, with actual context growth, tool calls delay and sub-agent spawning preserved, rather than on single-request benchmarks. NVIDIA Vera Rubin NVL72 agentic performance results are now live on the SemiAnalysis AgentX dashboard. On the DeepSeek V4 Pro model, Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72.Agentic workloads have a different shape from the request-response tasks that inference benchmarks have historically measured. A single session can accumulate hundreds of thousands of input tokens as agents reason, call tools and spawn sub-agents — roughly 15x the token volume of a simple chat request — with wide variability in input and output length across requests. Capturing that end to end takes a benchmark built around full agent trajectories. AgentX does that, complementing standardized suites like MLPerf Inference that measure per-request performance across a broad range of models and scenarios. The efficiency gain translates directly into AI factory economics. Up to 30x higher throughput per megawatt means up to 30x more agentic work from the same energy footprint, and the AgentX results show up to 45x lower cost per million tokens. In power-constrained deployments, throughput per megawatt determines how much revenue an AI factory can generate, and cost per million tokens determines the margin on it.This performance is the product of extreme codesign including the NVL72 scale-up domain and sixth-generation NVLink interconnect, NVFP4 precision on fifth-generation Tensor Cores, and an inference stack spanning NVIDIA TensorRT LLM and NVIDIA Dynamo. Vera Rubin is in full production and scaling across the ecosystem, and performance will continue to improve with ongoing software optimization.Learn more about the platform architecture behind these results. Tuesday, Sept. 15, 9:55 a.m. PT Startups Celebrate Performance Results With NVIDIA Vera CPUStartups are putting the NVIDIA Vera CPU to the test for running agent workloads and other demanding applications, celebrating the results. Perplexity benchmarked the Vera CPU for its new SPACE secure sandbox platform for agentic AI, showing 1.9x faster sandbox starts. Daytona put the Vera CPU through agentic workloads, citing “serious gains for agentic execution.” ClickHouse shared results for Vera CPU on ClickBench, its open benchmark for analytical databases. Vera was the fastest machine the team had measured so far. It’s a “strong signal of what’s ahead for CPU performance for data-intensive workloads,” ClickHouse said. DeepInfra ran a number of benchmarks on the Vera CPU, showing it winning on all metrics including 2.2x faster orchestration step latency.Prime Intellect found that the Vera CPU maintained high bandwidth and low, consistent memory latency as more workloads ran in parallel — the kind of predictable performance needed for agentic AI.Redpanda said its benchmark of the Vera CPU showed 5.5x lower latencies and 73% higher throughput than other CPUs.Starburst released benchmark results that showed the Vera CPU delivered 3x faster query throughput than other CPUs.Kinetica saw 2.7x faster analytical query performance with Vera CPUs compared with traditional CPUs. 22 Tuesday, Sept. 15, 9:55 a.m. PT NVIDIA NVLink 6 Keeps AI Factories Running at Massive ScaleAs AI factories scale to hundreds of thousands of GPUs, reliability becomes a performance feature. Training and inference workloads depend on continuous operation, making transient errors, signal degradation and hardware failures inevitable challenges. NVIDIA NVLink 6 addresses these challenges with a multilayer resiliency architecture designed to detect, contain and recover from faults before they affect applications.At the physical layer, custom forward error correction, physical layer retry and universal physical layer recovery help maintain a lossless fabric with minimal latency impact. At the network layer, credit-based flow control, dynamic routing and link rebalancing contain faults locally and prevent cascading stalls that can reduce AI factory throughput.Learn more about NVIDIA NVLink 6.