OpenAI's Jalapeno chip outperforms Nvidia's GB200 in OpenAI-supplied benchmarks
SemiAnalysis watched the tests in OpenAI's own lab using OpenAI's numbers, and calls the Nvidia comparison "somewhat incomplete and unfair."
Estimated reading time: 5 minutes
What happened
OpenAI published the first benchmark results for Jalapeno, an AI inference chip it has co-developed with Broadcom, on August 25, 2026, timed to the Hot Chips semiconductor conference, according to OpenAI’s company blog.
On SemiAnalysis’s InferenceX benchmark suite, Jalapeno delivered 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia’s GB200 and GB300 rack systems, SemiAnalysis reported. The comparison pits a 700W Jalapeno part against Nvidia parts rated at 1,200W to 1,400W. SemiAnalysis says it verified the runs in person at OpenAI’s lab, with OpenAI engineers present, and cross-checked the results against its own “chilisim” simulator, which it states is accurate to within 5%, per the same report.
All the benchmark numbers in that report were supplied to SemiAnalysis by OpenAI. SemiAnalysis did not run the full InferenceX suite itself, and it did not run its own AgentX benchmark — its preferred proxy for realistic production workloads — on Jalapeno at all, SemiAnalysis discloses. The publication also calls its own Nvidia comparison “somewhat incomplete and unfair,” arguing Jalapeno’s more appropriate rival is Nvidia’s next-generation Rubin platform, which is already shipping to customers while Jalapeno exists only as engineering samples, per SemiAnalysis.
The tested workloads were single-turn with an 8k context length only, with no long-context or multi-turn scenarios included. Most of the headline figures came from an early “A0” silicon stepping; a separate, less-tested “B0” stepping is claimed to add roughly 25% more performance per watt, SemiAnalysis notes. On a per-chip basis, Jalapeno is specified at 13.4 PFLOPs of MXFP4 compute at roughly 700W, against Rubin’s 17.5 PFLOPs at 900-1,150W; at rack level, Jalapeno systems are specified at 1.7 exaFLOPS of 4-bit compute with 27.5TB of HBM4 memory and near 2PB/s of memory bandwidth, according to SemiAnalysis.
Richard Ho, OpenAI’s head of hardware, said: “The bottom line is that the results show a very, very significant performance advance over state of the art. Jalapeno can serve more AI work per unit of power, while also returning responses more quickly,” Ho told TechCrunch.
The Register, reviewing the same disclosures, notes the Nvidia comparison used prior-generation (2024-2025) Nvidia systems and excluded speculative decoding, an optimization technique that improves real-world Nvidia inference performance, and cautions readers to “take all of these claims with a grain of salt,” The Register writes.
OpenAI plans to deploy Jalapeno in small volumes by the end of 2026, with meaningful scale-up in 2027; a second and third generation of the chip are already in development, per OpenAI.
What this means (and what it does not)
The results give OpenAI a public data point for a narrative of hardware independence from Nvidia, which supports its compute-cost and capital-expenditure story to investors and partners, as OpenAI’s own announcement frames it. Broadcom, Jalapeno’s co-designer and manufacturer, benefits from being positioned as a proven AI-silicon partner to a leading AI lab. SemiAnalysis, given exclusive in-person access, gains from being seen as the authoritative independent voice in AI hardware benchmarking, and its InferenceX suite gains prominence through the coverage — even as the firm itself flags the comparison’s limits, per SemiAnalysis.
What this does not mean is that Jalapeno has been shown to beat Nvidia’s current or next-generation hardware under independently controlled conditions. The benchmark numbers came from OpenAI, not from a suite SemiAnalysis initiated and ran itself. The comparison baseline, GB200/GB300, is not Nvidia’s newest shipping platform — that is Rubin, which SemiAnalysis itself says is the fairer comparator, per SemiAnalysis. Nvidia is the implicit target of the comparison and has not been shown to have responded. The Nvidia baseline also excluded speculative decoding, an optimization that would likely improve its real-world figures, per The Register.
What we still do not know
No benchmark run by a party independent of OpenAI’s invitation exists yet: SemiAnalysis watched tests in OpenAI’s own lab, on OpenAI-selected workloads, using OpenAI-supplied numbers, rather than running an end-to-end benchmark it controlled itself, per SemiAnalysis. Testing was limited to single-turn workloads at 8k context length, so there is no published data on latency or throughput at longer context lengths, in multi-turn conversations, or across varied batch sizes. No cost-per-inference-token or total-cost-of-ownership comparison has been published by any source reviewed here, so claims about the economics of AI deployment remain unsupported by cost data. Jalapeno is not in production: most published figures used early “A0” engineering samples, and the more efficient “B0” stepping is far less documented. SemiAnalysis did not run its own AgentX benchmark, its preferred proxy for realistic agentic workloads, on Jalapeno, so there is no data on its performance outside the narrower InferenceX conditions, per SemiAnalysis. It also remains unclear whether outside, uninvited third parties will be able to obtain Jalapeno hardware or reproduce these results independently before volume production begins in 2027.
Sources & Bylines
Every source cited in this article, gathered in one place.
- https://openai.com/index/jalapeno-first-results/
- https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
- https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/
- https://www.theregister.com/systems/2026/08/25/openais-upcoming-jalapeno-chip-looks-like-itll-be-an-inference-beast/5292052
Also available in Portugues (BR)