OpenAI Jalapeño: Better Than Nvidia Blackwell
Key Points:
- OpenAI has developed "Jalapeño," a custom AI inference chip built exclusively for large language model (LLM) inference, achieving a rapid ~16-month design-to-tapeout cycle starting mid-2024, in partnership with Broadcom.
- Jalapeño is a generalized inference chip optimized for performance per watt (perf/W) across various models and workloads, outperforming Nvidia, AMD, and Google's chips on multiple benchmarks without relying on specialized inference optimizations like Multi Token Prediction.
- Architecturally, Jalapeño uses HBM4 memory, a weight stationary systolic array, 64-bit scalar cores, and out-of-order execution to minimize latency and memory movement, enabling high efficiency on low-latency and high-throughput inference tasks; it also employs a unified prefill-decode design to maximize hardware utilization.
- The chip is supported by a sophisticated software stack including the Gluon kernel programming language built on Triton, with significant use of OpenAI’s Codex AI to write and optimize kernels rapidly, demonstrating fast software bring-up and iterative performance improvements.
- Jalapeño systems scale with a rack-level design featuring 128 ASICs per rack connected via high-bandwidth local and global networks, supporting up to 2,048 XPUs across 16 racks, with power efficiency and deployment scalability prioritized for large datacenter operations targeting 100MW scale.