OpenAI Benchmarks Jalapeño Against Nvidia's Racks
Ricardo Argüello, September 6, 2026
CEO & Founder
General summary
On August 25, 2026, the first results landed for Jalapeño, the inference chip OpenAI designed with Broadcom. By OpenAI's own measurements it delivers 1.5x to 1.9x the throughput of Nvidia's GB200 and GB300 racks, with end-to-end latency 1.7x to 3.6x lower and estimated power draw 40 to 60% lower. Initial availability is late 2026, volume production in 2027.
- Jalapeño is inference-only, designed by OpenAI and implemented by Broadcom, nine months from design to tape-out
- One system carries 128 accelerators, 1.7 exaFLOPS of 4-bit compute and 27.5 TB of HBM4 memory
- OpenAI reports 1.5x to 1.9x higher throughput than Nvidia's GB200 NVL72 and GB300 NVL72 racks
- It reports end-to-end latency 1.7x to 3.6x lower on GPT-OSS-120B, DeepSeek R1 and Kimi K2.5
- OpenAI supplied every figure and SemiAnalysis only witnessed the runs; initial availability is late 2026 with volume production in 2027
You have spent years shipping product with the only trucking fleet that exists, and that freight rate went into your price list as a fixed number. A competitor to the carrier shows up with trucks that haul twice the load on half the fuel. Your product does not change. What changes is which shipments made sense and which ones you crossed off as too expensive. That is what a more efficient inference chip does to your list of AI projects.
AI-generated summary
Nine months from design to tape-out.
That is the Jalapeño number I keep coming back to, and it is not the one making headlines. A chip of this size used to take two to three years. OpenAI says its own models helped compress the design cycle, and Broadcom did the silicon implementation.
The benchmarks are louder. On August 25, OpenAI published first results claiming 1.5x to 1.9x the throughput of Nvidia’s GB200 and GB300 racks, end-to-end latency 1.7x to 3.6x lower on GPT-OSS-120B, DeepSeek R1 and Kimi K2.5, and estimated power draw 40 to 60% lower.
Keep the label attached. OpenAI supplied every number. SemiAnalysis witnessed the InferenceX runs in its lab and confirmed the GSM8k evals landed on par with Nvidia, but it did not run the full suite or see AgentX results. The part is not shipping, and vendors publish the scenario where they win.
Your per-token price has a hardware floor
When you negotiate an API contract or compare model providers, you are looking at the far end of a chain. Under that price sits the real cost of running a model, and hardware sets it.
For three years that hardware was general-purpose. The same parts that train models also serve them, which leaves capability idle that inference never uses.
Jalapeño trains nothing. Richard Ho, OpenAI’s VP of hardware, described the goal as minimizing data movement and communication delay. One system carries 128 accelerators, 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 and 2 petabytes per second of memory bandwidth.
Specializing usually beats serving two masters. That has been the hardware story for forty years.
A shortening cycle beats a faster part
If nine months holds as the new cadence, the refresh cycle for AI hardware roughly halves. Whoever holds the best part today holds it for far less time than an enterprise infrastructure decision lasts.
Sign a three-year commitment against a cost reality that will turn over twice inside the term, and the commitment is the risk, not the chip.
What this actually asks you to do
Nothing involving hardware purchases. Something cheaper.
Your company has a list of shelved use cases. It lives in a 2025 email thread or a slide, and the reason written beside each one is some version of “too expensive per call.”
That list was built against a token price that has already moved and will move again when this class of silicon reaches volume in 2027.
We wrote about that trap in the per-token price lies, and about a concrete instance of it in Astra’s ten math proofs and the cost of inference. Same conclusion both times: a use case killed on cost is not dead, it is paused, and it deserves a calendar entry.
Concretely, this week
Find the list. If it does not exist, writing it is the first job: every case someone proposed and someone killed, with the reason beside it.
Split it in two. Cases killed on cost, and cases killed for anything else. The ones that died because the process was unclear or because nobody owned the outcome are not fixed by any chip, ever.
Put a date on the cost pile. Every six months, recompute at current prices. It takes half an hour and it is the highest-return review I know of, because the expensive part, understanding the process, is already done.
Jalapeño reaches volume in 2027. The list can be reviewed Tuesday.
Recompute one case you shelved on costFrequently Asked Questions
It is OpenAI's first custom chip, designed with Broadcom and dedicated entirely to inference, meaning running already-trained models. It was unveiled in June 2026 and its first performance results were published on August 25. Initial availability is late 2026, with volume production arriving in 2027.
OpenAI reports 1.5x to 1.9x higher throughput than Nvidia's GB200 NVL72 and GB300 NVL72 racks, end-to-end latency 1.7x to 3.6x lower on models including GPT-OSS-120B, DeepSeek R1 and Kimi K2.5, and estimated power consumption 40 to 60% lower. OpenAI supplied the figures. SemiAnalysis witnessed the InferenceX runs in its lab and confirmed the GSM8k evals came out on par with Nvidia chips, but it did not run the full suite or see AgentX results.
Because the per-token price you pay reflects what it costs to run those models on general-purpose hardware. When the hardware specializes, the floor under that price moves, and so does the list of use cases that were too expensive per call. Projects you shelved on cost deserve a scheduled re-review.
Training creates the model and needs large memory and sustained compute over weeks. Inference runs the finished model to answer queries and rewards low latency and energy efficiency per response. Jalapeño is built only for inference, unlike Nvidia and AMD parts that serve both workloads with one design.
Related Articles
NVIDIA Guaranteed $105B of OpenAI's Lease Obligations
NVIDIA backstopped up to $105B of OpenAI's Ohio lease. It only pays if OpenAI fails, and it terminates the moment OpenAI earns a credit rating.
OpenAI Astra: ten open problems for $2,000 in tokens
OpenAI says an internal version of Astra cracked ten long-open problems, each with a Lean 4 certificate. The tokens would cost about $2,000 at Sol API rates.
Perplexity at $30B on $750M, and Nvidia May Fund It
The Information reports Nvidia in talks to back Perplexity above $30B. In January it committed its entire annual revenue to three years of Azure compute.
Bristol Myers Squibb Is Building Its Own AI Factory, Not Renting It
BMS is the third pharma company in nine months to build its own AI supercomputer with NVIDIA. What it rents instead is a different layer entirely.
Pinecone Nexus Beat GPT-5.5 by One Point, at 77% Less
Coverage said Nexus outscored frontier models. The primary says 47.4% against 46.4%. What the knowledge layer actually bought was cost, on the cheaper model.
AI Prices Fell 40x. Most Marketing Teams Didn't Notice.
Fixed-capability AI inference dropped 40x per year since 2023. The value is not in this week's frontier model, it's in the workflows you never systematized.