NVIDIA Releases Groq 3 LPX, Long Context Inference Speed Breaks Through 3,400 Tokens per Second
NVIDIA has unveiled Groq 3 LPX, an interactive AI inference accelerator designed for the Vera Rubin platform, targeting agent reasoning scenarios characterized by long contexts, small batch sizes, and low latency. Third-party evaluation firm Artificial Analysis tested the chip using the Gemma 4 31B model and found that under a 100K input context, Groq 3 LPX achieved a median throughput of 3,431 tokens per second; at a 10K context length, it reached 3,382 tokens per second, demonstrating minimal performance degradation as context length increases. NVIDIA stated that generating 5,000 tokens takes approximately 1.5 seconds on this system, whereas systems delivering 100 tokens per second would require about 50 seconds. The Groq 3 LPX rack comprises 256 LP30 processing units with a total of 128 GB of SRAM, where each chip features 96 interconnect links operating at 112 Gbps. Its compiler schedules data transfers between chips before tasks begin execution and overlaps computation and communication in fine-grained detail, thereby reducing fixed communication overhead associated with tensor parallelism during small-batch inference. In SPEED-Bench programming tests, the system’s median throughput reached 4,767 tokens per second. NVIDIA plans to pair Groq 3 LPX with Vera Rubin NVL72, separating prefilling, decoding, attention, and FFN tasks to provide inference capabilities for long-context and ultra-large-scale multi-agent systems.
