OpenAI’s Jalapeño Benchmark Promises Faster, Cheaper AI

OpenAI’s Jalapeño Benchmark Promises Faster, Cheaper AI

OpenAI says its custom Jalapeño inference chip delivered up to 1.9 times greater performance per watt than the Nvidia Blackwell systems tested. Image: Jonathan Kemper/Unsplash

OpenAI says its Jalapeño chip delivered up to 1.9 times more performance per watt than Nvidia Blackwell systems in its first published benchmark tests.

Written By
Eric Mboizi
Eric Mboizi
Aug 27, 2026
We may earn from vendors via affiliate links or sponsorships. This might affect product placement on our site, but not the content of our reviews. See our Terms of Use for details.

In June 2026, Broadcom CEO Hock Tan and President Charlie Kawwas delivered Jalapeño to OpenAI CEO Sam Altman and President Greg Brockman. The custom AI inference chip was designed by OpenAI and co-developed with Broadcom.

OpenAI has now published its first benchmark results for Jalapeño following tests using InferenceX, a public AI inference benchmark from SemiAnalysis. The company said the chip delivered 1.5 to 1.9 times greater peak performance per watt than the Nvidia Blackwell systems used for comparison.

The results were produced by OpenAI rather than an independent testing organization, and the company normalized them using each accelerator’s published power rating.

What Jalapeño brings

Jalapeño was designed to work across different models, not only OpenAI’s systems. The ASIC (Application-Specific Integrated Circuit) was tested with GPT-OSS 120B and two non-OpenAI models: DeepSeek R1 670B and Kimi K2.5 1T. OpenAI said the results demonstrate that the architecture is not restricted to its own models.

AI inference has several phases with different bottlenecks. During prefill, the system processes the user’s prompt, which is compute-intensive. During decode, it generates the response token by token and relies more heavily on memory bandwidth. Communication between cores and chips can add further latency and reduce an AI model’s responsiveness.

Jalapeño has been designed to combine both high batch throughput and real-time responsiveness. OpenAI says that it has built a flexible AI accelerator “that can support changing model architectures, excel at both prefill and decode, and adapt as the balance between them changes, a defining feature of agentic workloads.”

One architectural feature behind this performance is a localized KV cache. OpenAI says model data can be explicitly placed and kept close to the required compute resources, reducing the time and power spent moving data during inference.

OpenAI says the resulting design can support high batch throughput and real-time responsiveness without making the same trade-off between throughput and latency found in some existing systems.

Advertisement

More must-read AI coverage

What Jalapeño means for enterprise users

Earlier this month, OpenAI launched a limited preview of an Ultrafast service tier that runs GPT-5.6 Sol at up to 14 times the speed of Standard processing. The service, powered by Cerebras, can generate up to 750 output tokens per second.

The Cerebras announcement came less than two months after OpenAI unveiled Jalapeño. Together, the announcements show that OpenAI is pursuing a multi-vendor infrastructure strategy rather than immediately replacing Nvidia and Cerebras hardware with its own chip.

OpenAI has confirmed that Jalapeño will complement rather than replace its partner-supplied accelerators. The company said: “Meeting growing demand for AI will require more compute from every available source. We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads.”

For enterprise customers, Jalapeño could eventually mean faster AI responses, greater service capacity and lower inference costs. OpenAI plans to begin deploying the chip within its own infrastructure by the end of 2026, but it has not said whether customers will be able to select the hardware directly or how the efficiency gains will affect API pricing. Those details will determine whether the benchmark produces a measurable advantage for businesses.

Read more: OpenAI’s first Jalapeño announcement explains five things businesses should know about the custom inference chip, including its potential cost savings and infrastructure trade-offs.

Eric Mboizi

Eric Mboizi is a technology news writer covering software development, emerging technologies, and the evolving digital landscape for TechRepublic and eWeek. He holds a bachelor’s degree in software engineering from Makerere University and has more than five years of experience creating technical content for developers and technology professionals. In addition to his work as a journalist, Eric is an Ethereum developer with more than four years of experience in blockchain technology. His hands-on development background gives him a practical perspective on software engineering, decentralized technologies, and the real-world implications of new technology trends.