OpenAI has introduced its custom inference chip, Jalapeno, which outperforms NVIDIA’s GB300 in speed and efficiency by combining high throughput with ultra-low latency through a novel hybrid memory architecture. This chip enables OpenAI to reduce inference costs, enhance user experience, and pursue a multi-chip strategy for scalable AI infrastructure, with plans for future generations of custom silicon to further optimize performance and affordability.
OpenAI has unveiled performance results for its first custom inference chip, Jalapeno, which reportedly outperforms NVIDIA’s GB300 system in both speed and efficiency. The benchmarks, conducted using the open-source InferenceX tool, demonstrate that Jalapeno achieves high throughput and ultra-low latency simultaneously—a first in the industry. This combination allows OpenAI to serve more users with faster response times, enhancing the overall user experience, particularly for agentic workloads like coding assistance.
Jalapeno’s architecture is distinct from traditional GPUs and TPUs, featuring a hybrid memory design that integrates both High Bandwidth Memory (HBM4) and SRAM. This design was developed from a “blank slate” approach, focusing on minimizing data movement, which is a major bottleneck in large language model inference. The chip’s novel architecture and programming model have enabled OpenAI to quickly optimize and deploy various AI models on the hardware, showcasing its flexibility and efficiency despite being a new design.
OpenAI’s decision to focus on inference rather than training with Jalapeno stems from the growing demand for inference compute power driven by increasing user activity. While NVIDIA remains a key partner for training hardware, OpenAI aims to diversify its inference infrastructure by integrating Jalapeno alongside other chips from partners like Cerebras. This multi-chip strategy allows OpenAI to optimize workloads based on cost and latency requirements, ultimately reducing inference costs and improving performance for end users.
Economically, Jalapeno is expected to significantly lower the cost per token for AI inference, with performance per watt improvements estimated between 1.8 to 4 times better than existing chips. This reduction in infrastructure costs aligns with OpenAI’s broader strategy to make AI services more affordable and scalable. The company is also addressing supply chain challenges, working closely with partners such as Broadcom and Celestica to scale production despite constraints in high bandwidth memory and wafer availability.
Looking ahead, OpenAI views Jalapeno as part of a long-term roadmap for custom silicon development, already working on second and third-generation designs. The company retains full ownership of the chip and system design, differentiating itself from competitors like NVIDIA, which controls a significant portion of server design content. If Jalapeno’s cost and performance benefits continue as projected, OpenAI anticipates a growing share of its inference workloads will run on its custom silicon, reinforcing its commitment to optimizing AI infrastructure from the hardware level up.