OpenAI reports its first custom inference chip, Jalapeño, can serve more AI work per unit of power while also returning responses more quickly. Jalapeño delivers both higher throughput and lower latency with one architecture, a feature often requiring tradeoffs in existing hardware systems. The chip achieves this efficiency across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T, delivering 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems.
For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance. OpenAI states these gains will help make increasingly capable AI more affordable and broadly available, demonstrating the chip’s versatility across models including GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T.
Jalapeño Achieves Industry-Leading AI Inference Performance
Jalapeño, OpenAI’s first custom inference chip, can serve more AI work per unit of power while also returning responses more quickly. This efficiency gain stems from a holistic design approach, integrating the chip itself with memory, networking, and serving software optimized for modern language models, particularly those powering interactive agents. OpenAI reports that this integrated design minimizes data movement and communication delays, allowing model state to remain local and processing units to operate more efficiently. Across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T, Jalapeño demonstrated broad compatibility across different architectures.
For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance. This focus on interactive workloads is particularly beneficial for applications where minimal delay is paramount. OpenAI’s ability to design the entire stack, from models to silicon, proved instrumental in Jalapeño’s development, allowing the team to leverage insights from real-world workloads to refine every layer. The chip’s architecture was specifically engineered to address the varying bottlenecks encountered during AI inference; prefill stages, which are computationally intensive, and decode stages, constrained by memory bandwidth, both receive optimized handling.
The company states that Jalapeño’s gains come from designing the chip, memory, network, software, and rack-scale system together around real language-model workloads, emphasizing the importance of this full-system approach. AI played a direct role in the chip’s creation, accelerating the design process and optimizing arithmetic circuits.
Engineers utilized AI tools to explore implementations, shorten design cycles, and continuously iterate on model workloads, completing the tapeout in just nine months. This AI-assisted design extends to programming the chip itself; using Codex with GPT‑Astra, the team rapidly brought three open-weight models to high performance within two months, demonstrating the architecture’s flexibility and the potential for AI-driven optimization. In some instances, AI-generated implementations of specific GPT‑OSS blocks ran 1.5 to 1.8 times faster than the existing human-expert-written implementations.
Increased efficiency translates to lower costs and broader accessibility of increasingly capable AI. The company plans to begin deploying Jalapeño within its compute infrastructure by the end of the year, marking the first step in a multigenerational platform with Gen 2 already in development and Gen 3 taking shape. OpenAI will continue to leverage accelerators from other vendors like NVIDIA while viewing Jalapeño as a key component in meeting the growing demand for AI compute and ensuring the benefits of artificial general intelligence are widely shared.
Tests using 1T revealed the chip’s versatility beyond OpenAI’s internally developed models, showcasing its ability to accelerate a range of AI workloads.
AI-Driven Design and Programming Speeds Chip Development
OpenAI is leveraging artificial intelligence not only to power its services, but also to fundamentally redesign how computer chips are created and programmed, a strategy demonstrated by the initial results of its custom inference chip, Jalapeño. The company’s approach extends beyond simply increasing processing speed; it encompasses a holistic integration of model design, product development, serving software, the chip itself, memory systems, networking, and rack-scale infrastructure.
This interconnectedness allowed engineers to apply insights gleaned from real-world workloads to optimize each layer of the system, resulting in a more efficient and responsive architecture. Jalapeño’s development showcased a unique synergy between AI and hardware engineering. The team utilized AI tools to accelerate the chip’s design process, shortening design, measurement, and verification loops, and continuously refining model workloads. This AI-assisted design extended to the arithmetic circuits within the chip, enabling the team to meet their performance targets on schedule.
OpenAI designed Jalapeño as a predictable programming target, facilitating both human and AI-driven optimization. Engineers define tasks using local tensors and explicit communication, while AI algorithms then optimize the mapping, placement, scheduling, and coordination of these tasks across the system. In specific instances, AI-generated implementations of GPT‑OSS attention and mixture-of-experts blocks ran 1.5 to 1.8 times faster than the existing human-expert-written implementations.
While these figures apply to selected blocks, they hint at a powerful new iterative development cycle. The company reflects its commitment to a fully integrated, AI-driven approach with the statement, “We used AI to design the chip, and designed the chip so AI could program it.” The architecture of Jalapeño prioritizes minimizing data movement and communication delays, a critical factor in maintaining performance. The system is designed to keep model state, including the KV cache, local to the processing units, while dynamically activating the optimal combination of compute, memory, and networking for each inference phase.
The network’s large domain ensures that the entire workload remains within a single connected system, further reducing latency and maximizing efficiency. By optimizing for both throughput and latency within a single architecture, Jalapeño addresses a common tradeoff in existing hardware systems, allowing for faster responses, more responsive agents, and more reliable access as demand for AI services grows. The company emphasizes that meeting growing demand for AI requires compute from every available source and will continue to utilize accelerators from vendors like NVIDIA alongside its own custom silicon.
See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
