On 700W vs 1,400W, what the SemiAnalysis InferenceX benchmark actually measures, and when this changes your API bill
OpenAI's Jalapeño has benchmarks now. The watt story is the real one.
Anti-AI
00
Skeptic
01
Neutral
00
Pro (practical)
02
Pro (hyped)
01
← Anti-AI · Pro-AI →
OpenAI gave Jalapeño its public debut at Hot Chips 2026 on August 25. The numbers they published are the first benchmark results using a third-party methodology — run against Nvidia's GB200 and GB300 rack systems on SemiAnalysis's public InferenceX benchmark.
They're legitimately good numbers.
On throughput per kilowatt, Jalapeño delivers 1.5 to 1.9 times more work per watt than the Nvidia systems it was measured against. On end-to-end latency, it's 1.7 to 3.6 times lower. On highly interactive workloads — the specific low-latency traffic pattern ChatGPT actually generates — it ran 2.1 to 4.1 times faster.
The chip draws about 700 watts. Nvidia's GB300 flagship rack draws roughly 1,400. That's the real headline, and it's a hardware fact, not a benchmark claim.
Source spread
- OpenAI — Jalapeño first results [hype] — OpenAI's own announcement. Leads with "industry-leading speed and efficiency." InferenceX tests cited; first-party runs on third-party methodology.
- Tom's Hardware — 700W ASIC vs 1,400W Nvidia flagship [builder] — Concrete on architecture and power draw. Notes Broadcom collaboration and AI-assisted design.
- Bloomberg — OpenAI Claims Chips Can Outperform Nvidia [skeptic] — "Claims" framing, appropriate. Notes deployment in small volumes end of 2026.
- SemiAnalysis — Jalapeño: Better Than Nvidia Blackwell [builder] — InferenceX benchmark methodology context. SemiAnalysis built the benchmark OpenAI ran.
Pros & cons
What checks out:
- The InferenceX benchmark is a public third-party methodology, not OpenAI's own eval. SemiAnalysis built it. OpenAI ran the tests on their hardware — that's a caveat — but the methodology is external, and SemiAnalysis has no obvious reason to flatter OpenAI.
- A 700W chip outperforming a 1,400W chip on throughput-per-watt is physically coherent. You're carrying half the thermal overhead for comparable compute.
- Tested against real production-grade models: DeepSeek R1, a 1-trillion-parameter version of Kimi K2.5, and OpenAI's open-source model. Not toy benchmarks.
- The 2.1–4.1x speedup on interactive workloads matters specifically for ChatGPT. Most inference benchmarks optimize for batch throughput, not conversational latency. Jalapeño was clearly designed for the latter.
What deserves scrutiny:
- OpenAI ran the hardware tests. SemiAnalysis provided the methodology. Those are different things — independent reproduction at scale doesn't exist yet because the chip isn't deployed.
- "In very small volumes" by end of 2026 is the timeline. Mass deployment is 2027. The capacity impact on OpenAI's actual infrastructure is minimal this year.
- The benchmark was tested on three models, all of which happen to be high-interest for OpenAI's own platform. Nobody ran Anthropic's or Google's workloads.
- OpenAI says their own models helped design Jalapeño. Useful capability demonstration, but it also means the chip was optimized for OpenAI's specific inference patterns — which may not generalize.
| Metric | Jalapeño (~700W) | Nvidia GB300 (~1,400W) |
|---|---|---|
| Throughput per kilowatt | 1.5–1.9× more | Baseline |
| End-to-end latency | 1.7–3.6× lower | Baseline |
| Interactive workload speed | 2.1–4.1× faster | Baseline |
| Power draw (approx.) | ~700W | ~1,400W |
| Mass deployment | 2027 | Now |
Samwise's take
What builders need to know
- Jalapeño doesn't change what you can do today. No builder-accessible Jalapeño hardware this year; mass deployment comes in 2027.
- Watch API pricing as the real signal. The first indication Jalapeño is in production will be a rate-card adjustment, not a product announcement.
- The interactive latency gains (2.1–4.1×) matter most for conversational and streaming applications — real-time voice, low-latency agent turn completion, streaming chat interfaces. If that's your workload, this is the metric to track.
- Treat the benchmark numbers as plausible, not independently verified. SemiAnalysis provided the methodology; OpenAI ran the tests. Independent reproduction doesn't exist yet.
- The AI-assisted chip design angle is worth watching as a capability signal, not a marketing claim. If it works, every AI company with enough scale will attempt the same vertical integration.
Further reading
- OpenAI — Jalapeño first results — official benchmark announcement
- Tom's Hardware — 700W ASIC vs 1,400W Nvidia flagship — power draw and architecture details
- Bloomberg — OpenAI Claims Chips Can Outperform Nvidia — straight news with appropriate skepticism
- SemiAnalysis — Jalapeño: Better Than Nvidia Blackwell — InferenceX benchmark methodology and context
- Axios — OpenAI says Jalapeño chip bests Nvidia — deployment timeline details
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.