Jalapeno Won the Watt and Tied on the Token. Only One of Those Reaches a Price Sheet.
OpenAI walked into Hot Chips on Tuesday, August 25, 2026 with the first published benchmarks for Jalapeno, the inference ASIC it built with Broadcom, and the headline is exactly as good as the headlines say. Against Nvidia GB200 and GB300 rack systems on SemiAnalysis's public InferenceX suite, Jalapeno delivered 1.5x to 1.9x more throughput per kilowatt at peak and 1.7x to 3.6x lower end-to-end latency. Dylan Patel summarized it in one line: usually first generation chips are not competitive, and this one is beating Blackwell and even Rubin.
I have been tracking this program since the tape-out. In June, when Jalapeno was unveiled, the number OpenAI put in front of everyone was roughly 50 percent lower cost per token than current Nvidia GPUs in early testing. That was the claim. Yesterday was the first time anyone put instruments on it.
So here is the sentence that matters, and it is not in most of the coverage. SemiAnalysis says the fair comparison is not Blackwell but Vera Rubin, because both parts use HBM4. Jalapeno still squeezes out more output tokens per megawatt than Vera Rubin. And on total cost of ownership per token, the two come out roughly even.
Roughly even. Not 50 percent lower. The efficiency claim survived contact with a benchmark. The cost claim, measured against the platform that will actually be in the racks next to it, did not.
What Was Actually Measured
Three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot's trillion-parameter Kimi K2.5. OpenAI supplied the numbers and SemiAnalysis verified some runs on-site in OpenAI's lab, which is better disclosure hygiene than most vendor benchmarks get and worth saying plainly. On GPT-OSS, Jalapeno hit roughly 1,400 tokens per second per user. On DeepSeek R1, it cleared 700 tokens per second on a single concurrent request.
| Spec | Jalapeno | Nvidia GB300 |
|---|---|---|
| Rated package power | 700W | 1,400W |
| Measured sustained | at or below 550W | not disclosed |
| All-in utility power | 1.18 kW | 2.55 kW |
| Memory | 216 GiB HBM4 | 288 GB HBM3E |
| Bandwidth | 15.4 TB/s | not compared |
| Trains models | no | yes |
| Shipping status | engineering samples | in customer racks |
Per watt of rated power, OpenAI's part carries roughly 50 percent more memory than the GB300. The Hot Chips deck says the bottleneck the architecture targets is exposing aggregate HBM bandwidth rather than adding more of it, which is a real design opinion and, on this evidence, a correct one.
Four Choices That Set the Size of the Win
None of these are cheating. All of them are normalization decisions that a vendor gets to make when it publishes its own numbers, and each one moves the gap.
| Choice | Effect on the lead |
|---|---|
| Normalize to published package TDP | Headline 1.5x to 1.9x. OpenAI's own appendix, using all-in utility power per accelerator, produces narrower gaps. |
| Compare against GB300, not Vera Rubin | Rubin is the HBM4 peer and the platform OpenAI itself agreed to deploy a gigawatt of in the second half of 2026. It was not in the test. |
| Single-token prediction on both sides | Against a GB300 running multi-token prediction, the peak efficiency lead shrinks to roughly 1.5x. Production Nvidia deployments commonly use MTP. |
| Pick the operating point | At the GB300's fastest previous time-between-tokens settings, OpenAI claims 8.6x to 104.3x. That is the number in the chart, and it is a latency-corner measurement, not a fleet average. |
Two of those cut the other way, and I want to be fair about it. Jalapeno posted its numbers without multi-token prediction and without speculative decoding while some comparison systems used them, so there is headroom left on the table. And a first-generation part landing anywhere near a mature platform is genuinely unusual.
Why the Tie Is the Story
Watts per token and dollars per token are different currencies, and only one of them ends up on an API price sheet. If you buy tokens, your bill is a function of total cost of ownership: silicon amortization, HBM, packaging, power, cooling, networking, and the utilization you actually achieve. Power is one line in that stack. Halving it is worth real money and it is not worth half the bill.
That is what the SemiAnalysis TCO note is telling you. Against Vera Rubin, the two parts land in the same neighborhood on cost per token. Which means that as of yesterday, nothing about the published inference price floor moved. No developer's cost per million tokens changed. What changed is who captures the margin between the cost of a token and the price of a token, and today that is a conversation between OpenAI and Nvidia, not between OpenAI and you.
The second-order effect is real, though, and it is the reason CNBC ran this as a margin story. Nvidia reports fiscal Q2 after the close today, with guidance around $91 billion against a record $75.2 billion data center quarter last time out. A first-generation ASIC from a customer that Nvidia is simultaneously financing, to the tune of up to $105 billion for the Ohio campus announced on August 17, is not a revenue threat this quarter. It is a price-negotiation artifact. OpenAI now has a published benchmark to put on the table.
Richard Ho, OpenAI's VP of hardware, said the quiet part out loud after the financing news: Nvidia is a really good partner, and we continue to need a lot of Nvidia. Sarah Friar framed Jalapeno as complementing the Nvidia, AMD, AWS, Cerebras, and CoreWeave relationships rather than replacing them. Both of those are true and both of them are also what you say when the leverage has shifted a few degrees and you would like to keep the shipment schedule.
The Claim I Think Is Underpriced
Buried under the perf-per-watt chart is the sentence with the longest half-life. SemiAnalysis wrote that the CUDA moat is potentially dead given how fast OpenAI can bring up new models on its silicon.
Consider what the bring-up actually looked like. Design started mid-2024. Final design went to fabrication in November 2025. Sixteen months end to end, nine of them from first chip design to finished blueprint. OpenAI used its own models in the loop: older generations on chip design, newer ones on programming and optimization. Then it ran three foreign models it did not train, including a trillion-parameter Kimi checkpoint, on first-silicon hardware well enough to publish.
The moat was never the instruction set. It was the years of kernel and compiler work that made a new accelerator painful to target. If a lab can compress that work using the models the accelerator exists to serve, the moat is not gone but it is measurably shallower, and the depth is now a function of model capability, which is the one variable in this industry that only moves one direction.
Three Counterarguments
Rubin ships, Jalapeno does not. Rubin systems are in customer hands. Jalapeno reportedly has not moved past engineering samples, with first deployment in OpenAI's own data centers targeted for later this year. A benchmark against a shipping platform from a part that is not yet in production is a promise with a chart attached.
The model list is stale. Nvidia and AMD have already published results on larger models, DeepSeek V4 Pro and Kimi K3 among them, that Jalapeno has not been tested against. Inference efficiency is workload-shaped. A part that wins on a 120B and a 670B does not automatically win on whatever ships in Q4.
HBM is the actual constraint. This is the one I would put money on. Samsung, SK hynix, and Micron have sold HBM capacity through 2027. Micron told the same conference on August 23 that HBM burns roughly three times the wafer area of DDR5 for equivalent capacity, and that the penalty widens each generation. SK hynix's CEO has called 2027 the worst year of the crunch. Scaling Jalapeno across the 10 gigawatt Broadcom agreement makes OpenAI a large new claimant on HBM4 supply that Nvidia currently dominates through multi-year allocation deals. Jalapeno sits in the same TSMC 3nm-class wafer queue, the same HBM queue, and the same advanced packaging queue as Blackwell and Rubin. You cannot design your way out of a line you are standing in.
Our Take
This is a real engineering result and a smaller economic one than the day's coverage implies, and the gap between those two facts is the entire piece. A first-generation inference ASIC beating a mature rack platform on throughput per kilowatt is the kind of thing that does not usually happen, and I do not want to talk anyone out of being impressed by it. But the June claim was 50 percent lower cost per token, the August measurement is roughly even against the right peer, and a claim that gets tested and comes back smaller is how you tell a benchmark from a press release.
The honest read: OpenAI bought itself a second source and a negotiating position, not a price cut. Vertical integration at this layer pays off in supply security and in what you can say to your largest vendor, and it pays off years later if generation two and three extend the curve. It does not show up in anyone's cost per million tokens in 2026.
Three things I am watching:
One, whether anyone publishes a Jalapeno versus Vera Rubin run on the same suite with MTP enabled on both sides. That is the only comparison that answers the question people think was answered yesterday.
Two, whether the second-generation part, reportedly approaching tapeout within months, arrives with an HBM4 allocation attached. Silicon without memory is a slide.
Three, whether any of this reaches a published price. We track the inference floor daily on our models tracker and cost calculator. If custom silicon is going to change what a token costs a developer, it has to show up there. So far it has not, and a tie on total cost of ownership is a fair explanation of why.
