LIVE
+++ Canada Hits Back With Tariffs of Up to 50 Percent +++ SELF: An Executable That Is Also a SQLite Database +++ Waymo Plans Driverless Taxis in Munich by Late 2027 +++ SpaceX Plans $100 Billion Spaceport in Louisiana +++ Suzuki e-Sky mini EV aims to beat BYD's Racco +++ Generalist raises $200M at a reported $3B valuation ++++++ Canada Hits Back With Tariffs of Up to 50 Percent +++ SELF: An Executable That Is Also a SQLite Database +++ Waymo Plans Driverless Taxis in Munich by Late 2027 +++ SpaceX Plans $100 Billion Spaceport in Louisiana +++ Suzuki e-Sky mini EV aims to beat BYD's Racco +++ Generalist raises $200M at a reported $3B valuation +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
DE EN
CHIPS

OpenAI's Jalapeño Chip Beats Nvidia in Inference Tests

OpenAI's first in-house accelerator posts more tokens per watt than Blackwell and Rubin on InferenceX — but only an early A0 sample was measured.

OpenAI's Jalapeño Chip Beats Nvidia in Inference Tests

Illustration: A data center aisle with two facing server racks and dense copper cabling under cool night lighting.

OpenAI's first in-house accelerator, Jalapeño, delivered more tokens per watt and lower response latency than Nvidia's Blackwell and Rubin systems in SemiAnalysis' InferenceX runs, though the silicon tested was a pre-production sample measured on a narrow workload.

At a glance

  • Built by Broadcom, fabricated by TSMC on N3P; CoWoS tape-out in November 2025.
  • Samsung HBM4 memory, 15.4 TB/s of bandwidth per package, 700-watt TDP per compute die.
  • InferenceX results: 1.5x to 1.9x more compute per watt and 1.7x to 3.6x lower end-to-end latency than shipping systems.
  • Only A0 silicon was measured, on 8k-context single-turn work; the AgentX suite was not run.
  • Volume ramp starts in 2027; late 2026 brings only very small quantities.

OpenAI has released the first performance figures for Jalapeño, the inference accelerator it designed in-house. On SemiAnalysis' public InferenceX suite, the part returned more tokens per second per user and more throughput per kilowatt than the best systems currently on the market. It is an inference chip only — it runs models, it does not train them.

What the numbers say

At peak throughput, SemiAnalysis reports 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end-to-end latency than commercially available systems. On interactive workloads, where responsiveness matters more than raw volume, the gap widens to between 2.1x and 4.1x.

Running DeepSeek R1 with a single concurrent request, Jalapeño cleared 700 tokens per second per user. On GPT-OSS 120B the testers logged roughly 1,400 tokens per second per user. With Kimi K2.5 held at an operating point of 100 tokens per second per user, SemiAnalysis puts the chip about nine times ahead of the next-best accelerator.

The caveat cuts in OpenAI's favor here. Jalapeño hit those figures without multi-token prediction and without speculative decoding, while the competing parts were measured in their best-performing configurations — which do use them.

The silicon

TSMC fabricates the chip on its N3P process, with Broadcom as design partner. Each compute die draws 700 watts and pairs with Samsung HBM4 delivering 15.4 TB/s per package. The A0 stepping was the one tested; a B0 stepping is already in the fab and is expected to add roughly 25 percent in performance per watt.

On paper, B0 reaches 13.4 PFLOPS of MXFP4 per compute die against 17.5 PFLOPS of dense NVFP4 for Nvidia's Rubin — close, at a lower power envelope. A single rack holds 128 Jalapeño chips, and 16 racks form a scale-up domain of 2,048 accelerators.

A paired compute-and-host system draws about 160 kW in total. Celestica handles the system-level design.

Why watts, not dollars

The focus on performance per watt follows from a constraint OpenAI states directly: it is limited by data center power, not by budget or floor space. Under that constraint, tokens per megawatt is the number that decides economics.

Richard Ho, who leads hardware at OpenAI, told TechCrunch the chip serves more AI work per unit of power while returning answers faster. The architecture gets there largely by refusing to move data — the KV cache stays local instead of crossing the network.

One design choice runs against industry fashion. Jalapeño does not split prefill and decode across separate chip pools; OpenAI argues that input-to-output ratios shift through the day, so fixed pools sit idle whenever the traffic mix changes.

What the benchmarks do not establish

SemiAnalysis is explicit about the limits. All the numbers came from OpenAI. Its team verified the InferenceX runs in person in the lab but did not run the full suite itself, and has not seen results from the harder AgentX suite.

Only single-turn workloads at 8k context were evaluated. Multi-turn, long-context traffic of the kind production systems actually carry was not tested, which leaves prefix caching and request routing unmeasured. Larger current models such as DeepSeek V4 Pro and Kimi K3 — the ones Nvidia and AMD are benchmarked against — do not yet run on Jalapeño.

SemiAnalysis also calls the Blackwell comparison incomplete, arguing Vera Rubin is the fair matchup because both use HBM4. On that footing, cost per output token comes out close to even. Rubin systems are already shipping to customers; Jalapeño exists only as engineering samples.

Timeline

Design work started in mid-2024 and taped out in November 2025 — about 16 months from the first hires to manufacturing. Three months of bring-up on real silicon produced the results published now.

Very small volumes are expected by the end of 2026, with the real ramp scheduled for 2027. The competitive baseline will likely have moved by then.

OpenAI used its own models during the design work. According to SemiAnalysis, that assistance cut SIMD area by 8 percent and matrix-engine area by 10 percent.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is OpenAI's Jalapeño chip?

Jalapeño is OpenAI's first custom ASIC, developed with Broadcom and fabricated by TSMC on N3P. It runs large language models for inference and does not train them.

Is Jalapeño faster than Nvidia's Blackwell and Rubin?

On SemiAnalysis' InferenceX runs it is, at 1.5x to 1.9x the performance per watt and 1.7x to 3.6x lower latency. But the test used a pre-production A0 sample on 8k-context single-turn work, and OpenAI supplied all the underlying data.

When will the Jalapeño chip be deployed?

Only engineering samples exist today. Very small volumes are planned for late 2026, with the production ramp beginning in 2027.