Luisa Crawford
Aug 24, 2026 17:11
NVIDIA’s Groq 3 LPX units a brand new commonplace in AI inference, delivering 3,431 tokens/second on a 100K context benchmark and redefining high-interactivity workloads.
NVIDIA’s Groq 3 LPX has redefined efficiency requirements for AI inference, reaching a world-class 3,431 tokens per second (TPS) on a 100K context benchmark, in accordance with an official weblog put up on August 24, 2026. Benchmarked by Synthetic Evaluation utilizing the Gemma 4 31B mannequin, this efficiency underscores Groq 3 LPX’s skill to deal with high-interactivity and long-context AI workloads on NVIDIA’s superior Vera Rubin platform.
The Groq 3 LPX is constructed round NVIDIA’s LP30 accelerators and rack-scale structure, providing 315 PFLOPS of FP8 inference compute and 128 GB of SRAM. The system’s concentrate on deterministic execution, low-latency token technology, and fine-grained scheduling allows it to excel in multiturn agentic classes and interactive workloads the place context grows considerably with every consumer interplay. For comparability, conventional fashions usually battle to keep up interactivity at this scale, particularly with lengthy enter contexts.
Synthetic Evaluation ran the benchmark with a 100K enter context size, measuring Groq 3 LPX’s skill to generate outputs at an unprecedented velocity. That is notably essential for agentic AI duties, comparable to coding or reasoning workflows, the place fashions should course of massive volumes of accrued context. NVIDIA experiences that such speeds translate to producing 5,000 tokens in simply 1.5 seconds, in comparison with 50 seconds at 100 TPS—an order-of-magnitude leap in effectivity.
Notably, the system additionally carried out effectively on a 10K context benchmark, delivering 3,382 TPS with minimal variation in latency. For coding duties, NVIDIA’s SPEED-Bench checks confirmed a median velocity of 4,767 output tokens per second, with 20% of duties exceeding 5,500 TPS, highlighting LPX’s versatility throughout completely different use circumstances.
Implications for AI Factories and Workloads
The Groq 3 LPX is positioned as a important element inside NVIDIA’s Vera Rubin platform, notably in AI factories the place high-interactivity serving tiers are important. Its skill to handle 100K+ context tokens whereas sustaining low latency may rework industries reliant on complicated, multiturn AI interactions, comparable to customer support automation, large-scale coding assistants, and real-time decision-making programs.
Key to this functionality is the LPX’s compiler-scheduled workload planning, which minimizes communication overhead between its 256 interconnected LPUs. By overlapping computation and communication at a fine-grained degree, the system achieves unmatched effectivity, even at small batch sizes the place conventional tensor parallelism methods usually falter.
NVIDIA’s Strategic Place
This announcement comes at a pivotal time for NVIDIA, which has cemented its management in AI {hardware}. Whereas the Groq 3 LPX will not be a standalone cryptocurrency-related product, its implications for AI-driven industries are immense. NVIDIA’s Vera Rubin platform, now enhanced by Groq 3 LPX, positions the corporate to dominate high-demand AI workloads, from real-time inferencing to generative AI purposes.
Shares of NVIDIA (NVDA) lately traded at $210.18, down 2.11% within the final 24 hours, amid broader market softness. Nevertheless, the corporate’s developments in AI inference expertise reinforce its long-term development narrative, notably in high-margin enterprise options. Latest rumors a few China-specific LPU product had been denied by NVIDIA, clarifying that no such roadmap exists, probably easing geopolitical issues for traders.
What’s Subsequent?
The Groq 3 LPX’s demonstrated efficiency opens the door to new AI purposes requiring each velocity and scale. NVIDIA has hinted at sustaining these speeds even with multi-hundred-thousand token contexts, which may additional revolutionize agentic AI use circumstances. Given its sturdy efficiency metrics and strategic integration with Vera Rubin, the Groq 3 LPX may set the benchmark for high-interactivity AI programs within the years to return.
Picture supply: Shutterstock








