OpenAI's Jalapeño chip: up to 1.9x more AI work per watt
OpenAI publishes first benchmarks for its Broadcom-built inference chip: more throughput and much lower latency — deployment starts late 2026.

Illustration · AI-generated (AI IN LIFE)
At a glance
- 1.5–1.9x more AI work per watt at peak throughput (OpenAI's measurements)
- 1.7–3.6x lower end-to-end latency versus comparison systems
- 2.1–4.1x higher performance on interactive workloads
- Tested on Kimi K2.5 (1 trillion parameters): max 550W against a 700W rating
- Developed in ~9 months; deployment from late 2026, broader rollout in 2027
OpenAI released the first benchmark results for Jalapeño on August 25, 2026 — the inference chip it co-developed with Broadcom. The headline numbers: 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. According to OpenAI, the architecture delivers both at once, where existing hardware typically has to trade throughput against response time.
What exactly was tested? OpenAI ran Jalapeño against a current Nvidia Blackwell system on SemiAnalysis's InferenceX benchmark, using open models such as GPT-OSS 120B, DeepSeek R1 670B and the one-trillion-parameter Kimi K2.5. On the largest model, measured sustained power draw stayed at or below 550 watts — the chip is rated for 700 watts.
For interactive workloads — chat and agent requests with many short response steps — OpenAI reports 2.1 to 4.1 times higher performance. Hardware chief Richard Ho told TechCrunch the results show "a very, very significant performance advance", saying the chip can "serve more AI work per unit of power, while also returning responses more quickly".
The development story is notable in itself: the reticle-sized ASIC went from concept to production in roughly nine months — unusually fast for a chip of this class. OpenAI used its own models in the design process; for selected compute blocks, AI-generated implementations ran 1.5 to 1.8 times faster than human-written code, the company says.
What happens next? The first Jalapeño systems are due to run in small volumes inside OpenAI's infrastructure by the end of 2026, with a broader rollout planned for 2027 — and a second generation is already in development. Ho also acknowledged that the competitive landscape may shift before full deployment: Nvidia, AMD and others will ship new inference hardware in the meantime.
FAQ
What exactly is Jalapeño?
Jalapeño is OpenAI's first custom inference chip, developed with Broadcom. It is optimized for serving model responses (inference), not for training.
What hardware was it compared against?
OpenAI benchmarked against a Nvidia Blackwell system on SemiAnalysis's InferenceX benchmark — Jalapeño led on throughput per kilowatt and tokens per user.
When does the chip go into service?
First systems are due in small volumes by the end of 2026, with the larger rollout in 2027. A second generation is already in development.


