OpenAI says Jalapeño chip delivers faster AI inference with less power

OpenAI's Jalapeño chip delivers 1.5-1.9x more AI work per watt and up to 3.6x lower latency than leading commercial systems, with deployment planned by year-end.

Categorized in: AI News IT and Development
Published on: Aug 26, 2026
OpenAI says Jalapeño chip delivers faster AI inference with less power

OpenAI has released benchmark results for Jalapeño, its first custom inference chip, showing the system delivers 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than leading commercial comparison systems. The company plans to begin deploying the chip in its compute infrastructure by the end of this year.

The results, measured on SemiAnalysis' InferenceX benchmark across three public models - GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T - position Jalapeño on the Pareto frontier for both performance per watt and latency. That means no comparison system achieved both better efficiency and lower latency simultaneously.

For customers, the practical effect is faster responses, more responsive agents, and more reliable access as demand grows. "Our mission is to ensure that artificial general intelligence benefits all of humanity. These gains will help make increasingly capable AI more affordable and more broadly available," the company said.

Design choices that avoid the throughput-latency tradeoff

Existing hardware systems typically force a choice between high throughput and low latency. Jalapeño was designed to avoid that compromise by treating language-model inference as a single workflow rather than separate phases. Prefill - processing the prompt - is compute-intensive, while decode - generating tokens one by one - is constrained by memory bandwidth. Communication between cores and chips can add further delays.

Jalapeño minimizes data movement by keeping model state, including the KV cache, local while activating the right combination of compute, memory, and networking for each phase. The network is integral to the architecture, with a large domain that lets an entire workload stay within one connected system. The result, the company says, is a "balanced and fungible accelerator" that can handle both prefill and decode efficiently and adapt as the balance between them changes - a defining feature of agentic workloads.

The chip is rated at 700 watts, though measured sustained power stayed at or below 550 watts on the tested workloads. The comparison systems were rated at 1,200 watts (GB200) and 1,400 watts (GB300).

AI helped design the chip, and the chip is built for AI to program

AI played a direct role in Jalapeño's development, helping the team move from initial design to tapeout in nine months. AI-assisted tools shortened design, measurement, and verification loops and optimized the chip's arithmetic circuits.

The architecture was also designed as a clear, predictable programming target. Engineers describe work through local tensors, explicit communication, and predictable synchronization. AI systems can then optimize how that work is mapped, placed, scheduled, and coordinated across the system.

Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months. For selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-expert-written implementations. Those figures apply to the selected blocks, not the full model.

Measured results across three models

On GPT-OSS 120B, Jalapeño delivered approximately 1.9 times higher peak mixed tokens per second per kilowatt (85,448 vs. 44,960) and 1.7 times lower end-to-end latency (1.03 seconds vs. 1.80 seconds) than the GB200 comparison system. Minimum time-between-tokens was 2.7 times lower at 0.69 ms versus 1.87 ms.

On DeepSeek R1 670B, the chip delivered 1.7 times higher peak mixed TPS/kW and 3.6 times lower end-to-end latency (1.65 seconds vs. 5.99 seconds) than the GB300 comparison system. Minimum TBT was 4.1 times lower at 1.43 ms versus 5.90 ms.

On Kimi K2.5 1T, the largest public model tested, Jalapeño delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency (1.56 seconds vs. 5.31 seconds). The company said its internal testing shows Jalapeño's advantage widens further on frontier OpenAI models, suggesting the architecture becomes more valuable as workloads grow larger.

At the comparison systems' best latency points, Jalapeño delivered dramatically more throughput: 53.7 times more on GPT-OSS, 104.3 times more on DeepSeek R1, and 56.1 times more on Kimi K2.5, measured at matched time-between-tokens.

Jalapeño expands what's possible for efficient inference

OpenAI describes the chip's impact in three tiers: ultra-fast-mode inference at efficiencies previously available only in fast mode; fast-mode inference at efficiencies previously available only in batched mode; and higher efficiency for batched-mode inference. The company says faster inference enables faster iteration and new use cases.

Jalapeño is the first generation of a multigenerational roadmap. Gen 2 is deep in development, and Gen 3 is taking shape. OpenAI said it will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads. The company is currently in production qualification, maturing software, preparing to operate Jalapeño at scale, and validating performance across more models.

For IT and development professionals, the significance is operational: faster inference at lower power cost changes what applications are feasible. Agentic workflows that require many sequential steps - where delays compound across an entire task - become more practical when time-between-tokens drops below 1 millisecond. For teams building on OpenAI's platform, the chip's deployment means more responsive applications without proportional cost increases. For teams running their own infrastructure, the benchmark methodology itself - performance per watt at matched user experience rather than raw chip specs - is worth adopting as a standard for evaluating inference hardware. Those working with Generative AI and LLM systems will find the throughput-per-watt gains directly relevant to cost modeling, while AI for IT & Development teams can expect faster model iteration and lower serving costs as the chip ramps into production.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)