What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble
Key Points:
- Nvidia showcased its Groq 3-based LPX racks delivering 3,400 tokens per second on Google's Gemma 4 31B model, making it 4x faster than the nearest competitor, Cerebras, which achieved 882 tok/s under similar conditions.
- Groq’s LPUs use SRAM-heavy dataflow architecture with extremely high memory bandwidth (150 TB/s) but limited on-die memory (500 MB per LPU), requiring Nvidia to distribute large models across multiple LPUs via Ethernet, with up to 256 LPUs per rack.
- Nvidia combines GPUs and Groq LPUs in a heterogeneous inference system, using GPUs for compute-heavy prefill phases and LPUs for memory-bandwidth intensive decode phases, significantly boosting inference speed and interactivity for AI agents.
- While the 31B parameter Gemma model fits within a single LPX rack, larger models like DeepSeek’s 671B parameter MoE require extensive hardware scaling, posing challenges for Nvidia’s architecture.
- Nvidia’s performance lead over Cerebras may be temporary as Cerebras is launching next-generation accelerators with doubled compute and memory bandwidth, and is partnering with AMD and AWS on hybrid GPU-accelerator systems, potentially narrowing the performance gap.