OpenAI introduced Jalapeno, their first custom AI ASIC developed with Broadcom, which significantly outperforms Nvidia GPUs in latency and energy efficiency for large-scale model inference by leveraging a specialized architecture optimized for their workloads. The chip, integrated into a scalable system, showcases AI-assisted design innovations and marks a strategic move towards custom silicon, with plans for future iterations while maintaining hardware flexibility.
OpenAI recently unveiled Jalapeno, their first custom-built ASIC processor designed specifically for machine learning and AI workloads, developed in partnership with Broadcom in just nine months. The chip is intended to power OpenAI’s inference tasks and is integrated into a large-scale rack system spanning two racks with numerous chips. The presentation highlighted the chip’s architecture, performance benchmarks, and the innovative use of AI in the chip design process. Jalapeno features a central compute die surrounded by six HBM memory modules and an IO die, optimized for OpenAI’s specific model workloads.
The development timeline for Jalapeno began in October 2024 with architecture design, followed by RTL coding, and culminated in tape-out and first silicon by November 2025. By mid-2026, the chip was running OpenAI’s code and supporting ChatGPT workloads. OpenAI emphasized two key performance metrics: latency (time to last token) and energy efficiency (tokens per joule), focusing on system-level performance rather than just chip throughput or time to first token. They benchmarked Jalapeno against Nvidia’s GPUs using open-source models and the Inference X benchmark suite to provide a transparent comparison.
Benchmark results showed Jalapeno outperforming Nvidia’s GPUs significantly in terms of tokens per second per watt and end-to-end latency across various model sizes, including GPT OSS 120B, DeepSeaCar 670B, and the larger Kim K2.5 trillion parameter model. Jalapeno demonstrated up to 100 times better performance per watt and up to 4 times lower latency, even when Nvidia used multi-token speculative decoding techniques. This performance advantage is attributed to Jalapeno’s dedicated AI architecture, optimized for OpenAI’s workloads, rather than repurposed GPU hardware.
Architecturally, Jalapeno is a spatial design running gluon programs with each core containing tensor, SIMD, and scalar engines alongside fast local L1 memory. The chip uses a collective network optimized for bandwidth and latency tailored to OpenAI’s model communication patterns. The system scales up to 128 chips interconnected via Broadcom Tomahawk 6 switches, enabling high bandwidth and low latency for tensor and expert parallelism. AI-assisted chip design tools helped improve unit areas beyond human-optimized baselines, accelerating development and enhancing efficiency.
Looking ahead, OpenAI plans to iterate on Jalapeno with second and third-generation chips, continuing to refine performance and efficiency. While Jalapeno is designed for OpenAI’s internal use, the company will maintain flexibility by continuing to use Nvidia and other vendors’ hardware to support diverse models and workloads. The chip’s success underscores the trend of hyperscalers building custom silicon tailored to their specific AI needs, balancing rapid innovation with the challenges of predicting future model requirements and market pricing.