Part #/ Keyword
All Products

NVIDIA Groq 3 LPX Enters Full Production

2026-08-26 13:48:36Mr.Ming
twitter photos
twitter photos
twitter photos
NVIDIA Groq 3 LPX Enters Full Production

According to NVIDIA, its Groq 3 LPX interactive AI inference acceleration system has entered full production following its debut at Hot Chips 2026 on August 24. Designed as an extension of the Vera Rubin platform, Groq 3 LPX features a single-rack configuration with 256 Groq 3 LPU accelerators, delivering ultra-fast token generation for high-response AI agent applications.

As AI development shifts from model training toward large-scale inference, demand for faster and more efficient AI reasoning is rapidly increasing, especially with the rise of agentic AI systems. NVIDIA explains that AI inference includes two key stages: Context processing, which handles long prompts, code, and accumulated agent states, and Generation, which produces the tokens required for users or AI agents to take action.

Groq 3 LPX is optimized for long-context processing and low-latency inference workloads. Each LPX rack integrates 256 LPU accelerators, 128GB of on-chip SRAM, 640TB/s scale-up bandwidth, and NVIDIA MGX rack architecture with liquid cooling. The system is designed to significantly improve token generation performance for real-time AI applications.

According to Artificial Analysis testing, Groq 3 LPX achieved a median output speed of 3,431 tokens per second with the Gemma 4 31B model under a 100K context workload, and 3,382 tokens per second under a 10K context workload, setting a new performance record for the model. NVIDIA’s SPEED-Bench evaluation also showed median token output speeds of 4,767 tokens per second.

NVIDIA stated that Groq 3 LPX can reduce complex agent tasks, such as coding workflows, from hours to minutes. Compared with competing platforms, it can provide up to four times faster response performance for latency-sensitive AI workloads.

Rather than replacing Vera Rubin NVL72, Groq 3 LPX is designed to complement it through a specialized inference architecture. Vera Rubin NVL72 handles the Prefill stage, processing context and building KV Cache, while Groq 3 LPX focuses on the Decode stage to accelerate token generation. This separation allows different hardware platforms to optimize their strengths for various AI workloads.

“NVIDIA Groq 3 LPX pushes the performance frontier of ultra-fast token generation for the agentic AI era,” NVIDIA CEO Jensen Huang said. “It enables a new generation of AI factories with higher throughput, efficiency, and responsiveness.”

AI cloud service provider Nebius will become the first company to deploy Groq 3 LPX through its Nebius Token Factory inference platform, allowing developers to access faster token generation for real-time AI agent applications.


* Solemnly declare: The copyright of this article belongs to the original author. The reprinted article is only for the purpose of disseminating more information. If the author's information is marked incorrectly, please contact us to modify or delete it as soon as possible. Thank you for your attention!