
According to NVIDIA, its Groq 3 LPX interactive AI inference acceleration system has entered full production following its debut at Hot Chips 2026 on August 24. Designed as an extension of the Vera Rubin platform, Groq 3 LPX features a single-rack configuration with 256 Groq 3 LPU accelerators, delivering ultra-fast token generation for high-response AI agent applications.
As AI development shifts from model training toward large-scale inference, demand for faster and more efficient AI reasoning is rapidly increasing, especially with the rise of agentic AI systems. NVIDIA explains that AI inference includes two key stages: Context processing, which handles long prompts, code, and accumulated agent states, and Generation, which produces the tokens required for users or AI agents to take action.
Groq 3 LPX is optimized for long-context processing and low-latency inference workloads. Each LPX rack integrates 256 LPU accelerators, 128GB of on-chip SRAM, 640TB/s scale-up bandwidth, and NVIDIA MGX rack architecture with liquid cooling. The system is designed to significantly improve token generation performance for real-time AI applications.
According to Artificial Analysis testing, Groq 3 LPX achieved a median output speed of 3,431 tokens per second with the Gemma 4 31B model under a 100K context workload, and 3,382 tokens per second under a 10K context workload, setting a new performance record for the model. NVIDIA’s SPEED-Bench evaluation also showed median token output speeds of 4,767 tokens per second.
NVIDIA stated that Groq 3 LPX can reduce complex agent tasks, such as coding workflows, from hours to minutes. Compared with competing platforms, it can provide up to four times faster response performance for latency-sensitive AI workloads.
Rather than replacing Vera Rubin NVL72, Groq 3 LPX is designed to complement it through a specialized inference architecture. Vera Rubin NVL72 handles the Prefill stage, processing context and building KV Cache, while Groq 3 LPX focuses on the Decode stage to accelerate token generation. This separation allows different hardware platforms to optimize their strengths for various AI workloads.
“NVIDIA Groq 3 LPX pushes the performance frontier of ultra-fast token generation for the agentic AI era,” NVIDIA CEO Jensen Huang said. “It enables a new generation of AI factories with higher throughput, efficiency, and responsiveness.”
AI cloud service provider Nebius will become the first company to deploy Groq 3 LPX through its Nebius Token Factory inference platform, allowing developers to access faster token generation for real-time AI agent applications.