OpenAI has announced Jalapeno, its first custom AI inference chip, developed in partnership with Broadcom, marking a significant shift in the AI hardware landscape. According to OpenAI and independent analysts, Jalapeno outperforms Nvidia’s latest GPUs and Google’s TPUs in key inference tasks, boasting superior speed, efficiency, and latency.
Breakthrough Architecture and Performance
Jalapeno is an application-specific integrated circuit (ASIC) designed from the ground up for large language model (LLM) inference. Unlike general-purpose GPUs, Jalapeno features a novel hybrid memory architecture that combines High Bandwidth Memory (HBM4) and SRAM, minimizing data movement—a major bottleneck in AI inference. This design enables high throughput and ultra-low latency, allowing OpenAI to serve more users with faster response times, particularly for demanding workloads like coding assistance.
Benchmarks conducted with the open-source InferenceX tool and observed by independent analysts at SemiAnalysis indicate that Jalapeno achieves up to 1.9x higher throughput per kilowatt and 3.6x lower latency compared to Nvidia’s GB300 GPU. In single-token prediction scenarios, Jalapeno reportedly delivers over 700 tokens per second per user on the DeepSeek R1 model, outperforming competitors even without speculative decoding or other optimizations. The chip also demonstrated strong results on other models, such as Kimi-K2.5 and GPT-OSS, with performance on par with or exceeding Nvidia hardware.
AI-Driven Chip and Software Design
A key innovation in Jalapeno’s development was OpenAI’s use of AI to accelerate both hardware and software design. The chip went from initial team hiring to manufacturing tape-out in approximately 16 months, an unusually rapid cycle for ASICs. OpenAI leveraged its Codex AI model to generate highly optimized kernel code, reducing reliance on scarce human experts and enabling quick adaptation to new AI architectures.
Strategic Implications and Industry Impact
OpenAI’s move to custom silicon is seen as a direct challenge to Nvidia’s dominance in AI hardware. By controlling its own inference infrastructure, OpenAI aims to reduce costs, improve scalability, and accelerate innovation. The company plans to integrate Jalapeno alongside chips from partners like Cerebras, optimizing workloads based on cost and latency requirements.
While OpenAI continues to rely on Nvidia for training hardware, Jalapeno is focused on inference, where demand is rapidly growing. The company is already working on second and third-generation chips and is addressing supply chain challenges to scale production.
Caveats and Future Outlook
Analysts caution that while Jalapeno’s initial results are impressive, most benchmarks were conducted in collaboration with OpenAI, and broader independent testing is needed. Some industry observers note that comparisons with Nvidia’s Blackwell GPU may not be entirely apples-to-apples, and real-world production workloads could reveal additional nuances.
Nevertheless, Jalapeno’s debut signals a new era of AI hardware competition, with OpenAI leveraging its own AI models to design chips and software, potentially accelerating the pace of AI innovation beyond what traditional hardware providers can match.
Sources
Internal sources
- OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing
- First Benchmark Results on an OpenAI GPU
- OpenAI BROKE the Industry Overnight
