OpenAI Unveils First In-House Inference Chip Results, Next-Gen in Production
On August 25, OpenAI disclosed initial performance benchmarks for Jalapeño, its first self-developed AI inference chip, revealing superior energy efficiency and throughput compared to Nvidia's GB200 and GB300 systems. OpenAI led the chip's architecture design, collaborating with Broadcom on implementation and networking, and Celestica for integration. The project progressed from initial design to tape-out in approximately nine months. Performance evaluations using SemiAnalysis's InferenceX benchmarking tool across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T models indicate that Jalapeño delivers peak throughput per watt ranging from 1.5 to 1.9 times higher than the comparison hardware while reducing generation latency. Detailed metrics demonstrate significant advantages across diverse workloads. In testing against the Nvidia GB200, Jalapeño achieved 1.9 times the peak throughput for GPT-OSS 120B, processing roughly 85,400 tokens per second per kilowatt versus 45,000, with a minimum inter-token latency of 0.69 milliseconds supporting user generation speeds up to approximately 1,459 tokens per second. For larger models including DeepSeek R1 and Kimi K2.5, Jalapeño outperformed the Nvidia GB300 with 1.7 to 1.9 times higher throughput efficiency. Under constrained latency conditions, the chip exhibited exceptional density; when maintaining identical single-user generation speeds, Jalapeño supported up to 104 times the concurrent task throughput of the GB300 for DeepSeek R1. The chip operates within a 700-watt rating, with observed power consumption remaining under 550 watts during tests, contrasting with the 1,200 to 1,400 watts calculated for Nvidia reference systems. Jalapeño employs a unified accelerator architecture to manage both prefill and decode phases within a single chip. OpenAI accelerated software development by leveraging AI tools such as Codex and GPT-Astra, completing model porting and optimization in two months. AI-generated kernels improved performance by 1.5 to 1.8 times in specific attention and MoE modules. While current results stem from the A0 engineering sample tested in short-context scenarios, the successor B0 version has entered manufacturing. The B0 iteration utilizes TSMC's N3P process and MXFP4 precision to achieve 13.4 petaflops of theoretical compute, integrating HBM4 memory with 15.4 terabytes per second of bandwidth and a dedicated I/O chip for scalable inter-rack communication. OpenAI targets deployment of Jalapeño within its computing infrastructure by the end of 2026, with second-generation development ongoing and third-generation planning already initiated. The company confirmed it will continue utilizing Nvidia and other partner chips to support varied operational needs. SemiAnalysis verified the InferenceX execution environment on-site, though performance data was supplied by OpenAI rather than independently measured. Current public results cover single-turn inference with limited context, and benchmarks for longer contexts or multi-turn agent workloads remain pending.
