On August 25, OpenAI published the first benchmark results for Jalapeño, its in-house inference chip (official announcement; SemiAnalysis, TechCrunch, and The Register covered it the same day, and SemiAnalysis is the only one of the three to report any on-site verification, more on that below). The numbers are striking. Against Nvidia’s shipping flagship Blackwell systems, Jalapeño does 1.5x to 1.9x more inference work per watt at peak throughput, with 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x faster interactive workloads. SemiAnalysis’s verdict is blunt: first-generation chips are usually uncompetitive, and OpenAI is beating Blackwell, even Rubin.

Silicon-industry common sense sits on the opposite side of these numbers: first-gen custom chips usually aren’t competitive, which is exactly the first half of SemiAnalysis’s line. So this report card allows two explanations. Either OpenAI broke the pattern, or the exam itself needs some qualification. Having read the material, my conclusion is both.

Who graded the exam

Start with the test. The benchmark is SemiAnalysis’s public InferenceX suite, which OpenAI’s announcement describes as measuring the full life of an AI request (from arrival to the last token out) rather than any single point metric, across three open-weight models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5. That sounds third-party enough, but SemiAnalysis is explicit in its own write-up: all numbers were provided by OpenAI. SemiAnalysis verified some runs in person at OpenAI’s lab and did not independently run the full InferenceX suite. Think of a student taking the exam at home while a proctor drops by to spot-check a few worked solutions. Spot checks beat nothing. They are not an independent retest.

The scope is narrow too. Only a single 8k-input, 1k-output workload has been tested so far, and AgentX, the part of the suite built to measure long-context, multi-turn agent tasks, has no public results at all. Production inference is messier than 8k/1k: multi-turn conversations, request routing, cache management, none of it on this exam.

The comparison target deserves a discount as well. Blackwell is a shipping product announced in March 2024, while Jalapeño exists as engineering samples (early silicon back from the fab for debug and validation, not production parts). Per TechCrunch, tiny deployments start in late 2026 and volume comes in 2027. By then it will face Nvidia’s next-generation Rubin platform, whose first gigawatt is slated for the second half of 2026. Richard Ho, OpenAI’s head of hardware, acknowledged as much to TechCrunch: by the time Jalapeño reaches full deployment, the competition may have advanced significantly.

Even after those discounts, what remains is real. SemiAnalysis notes the fairer comparison is Rubin, which also uses HBM4, and by the data it saw on site, Jalapeño still leads on tokens per megawatt. The competing systems ran with multi-token prediction enabled (a software speedup that has the model guess several tokens at once and verify them quickly); Jalapeño ran without it. Ho’s words: “The bottom line is that the results show a very, very significant performance advance.”

Why a first-gen chip can win: decode wants bandwidth, not FLOPs

Inference splits into two phases, and OpenAI’s announcement discusses them separately. Prefill, reading in the prompt, is compute-hungry like training. Decode, emitting tokens one at a time, is bound by memory bandwidth: each generated token typically requires streaming model weights from memory through the compute units (all of them for a dense model, the activated experts for MoE), so the arithmetic units spend most of their time waiting on data. The decode bottleneck is bandwidth, not FLOPs.

Jalapeño’s specs, as disclosed by SemiAnalysis and The Register, read like a bet on exactly that: TSMC N3P process, 216GB of HBM4 per package at 15.4TB/s of bandwidth, plus a large on-die SRAM cache. OpenAI says model state, including the KV cache (the intermediate results of the conversation context, reread on every generated token in long chats and agent runs), can be explicitly placed and kept local to the chip. The stated design goal was minimizing data movement and communication latency. Raw compute, by contrast, is not the story: the 13.4 PFLOPS of 4-bit compute SemiAnalysis quotes belongs to the upcoming B0 stepping, not the A0 silicon that produced these benchmarks, and The Register ran the numbers on competing rack systems, which deliver 1.46x to 2x more raw compute. This chip wasn’t built to out-FLOP anyone. The entire bet is on moving data.

A first-gen design can afford a trade-off this extreme because it answers only for inference. It doesn’t have to carry training plus a zoo of customer workloads the way a general-purpose GPU does. One tempting misreading needs correcting, though: SemiAnalysis specifically pushes back on the claim that Jalapeño can only run OpenAI’s own models. It is a general inference chip capable of running many models and workloads, and all three models in this public test are external open-weight models. What is genuinely proprietary is the software. OpenAI’s announcement says its frontier models participated in the model-to-chip mapping, scheduling, and low-level kernel work, and that AI-generated implementations of selected GPT-OSS attention and MoE modules run 1.5x to 1.8x faster than versions written by expert engineers. SemiAnalysis’s read is worth relaying: if Jalapeño succeeds, the industry’s attachment to general programming models and general compilers is being displaced by frontier models writing the low-level software themselves. Whether the CUDA barrier actually erodes as a result is, so far, an analytic judgment from SemiAnalysis, not something that has happened.

The real ledger: cost sheets and negotiating tables

The unit economics of inference are roughly: cost per million tokens ≈ (chip depreciation + power and cooling + datacenter amortization) ÷ token output. SemiAnalysis says OpenAI’s binding constraint right now is datacenter power, not budget. Power is the hard limit, so performance per watt directly sets how many tokens the same building can produce. A 1.5x to 1.9x per-watt gain, on this test’s terms, means at least half again more tokens from the same electricity.

The other half of the ledger is purchase price. Nvidia’s company-wide gross margin has run between 70 and 75 percent in recent years (FY2025 full-year GAAP gross margin: 75%). That is a company-level figure, not the margin on any particular accelerator, but it establishes that a GPU’s price carries a sizable profit layer. The custom-silicon route bypasses exactly that layer: Broadcom takes its own cut, which SemiAnalysis believes runs below Nvidia’s level, and part of the total-cost-of-ownership advantage comes from that margin gap. How much OpenAI actually saves, it hasn’t disclosed. Performance per watt and purchase price are two separate levers, and only together do they make the full arithmetic of OpenAI’s chip program.

Volume is still years away, though. Jalapeño ships in small quantities at year-end, and the 10-gigawatt Broadcom program runs through the end of 2029. The Nvidia side of the ledger has already been rewritten once: the September 2025 letter of intent for at least 10 gigawatts of Nvidia systems, under which Nvidia planned to invest up to $100 billion progressively as deployments came online, was restructured in February 2026. Per Financial Times reporting, it became a direct equity stake of roughly $30 billion in OpenAI’s latest fundraising round, and OpenAI is expected to spend much of that fresh capital on Nvidia chips. Add the 6-gigawatt agreement with AMD (with warrants for up to 160 million AMD shares), and my read is that in the near term Jalapeño changes OpenAI’s negotiating position more than its procurement mix. A custom chip that a third party has spot-checked on site, and that pencils out on paper, is the hardest external option you can bring to a procurement negotiation.

Google walked a similar road with the TPU. When it was made public in 2016 it had already been running in Google’s datacenters for over a year. Google Cloud still adopts each new Nvidia generation (including Blackwell GB200 NVL72 instances), but by Google’s own account the TPU is the backbone of AI workloads across nearly all its products. Google doesn’t disclose how it negotiates GPU purchases, but it is hard to imagine those orders arriving at the table without the in-house alternative behind them; that part is my inference. One difference: Google has rented TPUs to cloud customers since 2017, while OpenAI has so far announced deployment only inside its own infrastructure. OpenAI frames the payoff as response speed, cost, and operational flexibility; bargaining power is the most immediate of those returns.

There is also a direct read for agent products. The 2.1x to 4.1x interactive speedup, and DeepSeek R1 running above 700 tokens per second for a single user at concurrency 1, land exactly on agents’ sore spot: agent tasks are long serial chains where each step’s generation must finish before the next begins, and OpenAI itself notes that latency accumulates along the steps. Faster output shrinks the generation share of the chain proportionally. Tool calls and network waits don’t speed up with it, so the whole chain won’t accelerate by the same multiple, but response feel and per-task cost improve for real. The irony is that AgentX, the benchmark built for long-context multi-turn agent workloads, is precisely the one with no published results. The scenario the pitch benefits most from is the scenario with the least data. I couldn’t find supplementary numbers anywhere I looked; this waits on OpenAI or SemiAnalysis.

What to watch

Three signals worth tracking. First, when AgentX and long-context numbers appear; until then, treat the agent-side gains as unproven. Second, 2027 volume and yield: SemiAnalysis says an improved B0 stepping is already at the fab with roughly 25% better performance per watt expected, but the climb from engineering samples to mass production still has yield and capacity-ramp gates ahead. Third, pricing: if inference costs really fall, does the price tag follow? My guess is the savings show up first in the margins of OpenAI’s own agent products and reach API prices later. That is speculation; I haven’t seen a public pricing plan to point to.

For Nvidia, the near-term revenue is not the problem: even the restructured deal has OpenAI expected to spend much of Nvidia’s $30 billion on Nvidia chips, and what needs buying still gets bought. My judgment is that what these benchmarks shake is a pricing assumption: a gross margin around seventy percent is easiest to defend while customers hold no credible alternative. When a customer of OpenAI’s scale produces a chip that a third party spot-checked on site, and that leads on performance per watt by the numbers OpenAI supplied, that margin stops being priced by technology and starts being priced by negotiation. You can discount the benchmark. You can’t discount that.

References