OpenAI's Jalapeño has benchmarks now. The watt story is the real one.
Jump to a section
See the story at a glance
OpenAI gave Jalapeño its public debut at Hot Chips 2026 on August 25. The numbers they published are the first benchmark results using a third-party methodology — run against Nvidia's GB200 and GB300 rack systems on SemiAnalysis's public InferenceX benchmark.
They're legitimately good numbers.
On throughput per kilowatt, Jalapeño delivers 1.5 to 1.9 times more work per watt than the Nvidia systems it was measured against. On end-to-end latency, it's 1.7 to 3.6 times lower. On highly interactive workloads — the specific low-latency traffic pattern ChatGPT actually generates — it ran 2.1 to 4.1 times faster.
The chip draws about 700 watts. Nvidia's GB300 flagship rack draws roughly 1,400. That's the real headline, and it's a hardware fact, not a benchmark claim.
Source spread
- OpenAI — Jalapeño first results [hype] — OpenAI's own announcement. Leads with "industry-leading speed and efficiency." InferenceX tests cited; first-party runs on third-party methodology.
- Tom's Hardware — 700W ASIC vs 1,400W Nvidia flagship [builder] — Concrete on architecture and power draw. Notes Broadcom collaboration and AI-assisted design.
- Bloomberg — OpenAI Claims Chips Can Outperform Nvidia [skeptic] — "Claims" framing, appropriate. Notes deployment in small volumes end of 2026.
- SemiAnalysis — Jalapeño: Better Than Nvidia Blackwell [builder] — InferenceX benchmark methodology context. SemiAnalysis built the benchmark OpenAI ran.
Pros & cons
What checks out:
- The InferenceX benchmark is a public third-party methodology, not OpenAI's own eval. SemiAnalysis built it. OpenAI ran the tests on their hardware — that's a caveat — but the methodology is external, and SemiAnalysis has no obvious reason to flatter OpenAI.
- A 700W chip outperforming a 1,400W chip on throughput-per-watt is physically coherent. You're carrying half the thermal overhead for comparable compute.
- Tested against real production-grade models: DeepSeek R1, a 1-trillion-parameter version of Kimi K2.5, and OpenAI's open-source model. Not toy benchmarks.
- The 2.1–4.1x speedup on interactive workloads matters specifically for ChatGPT. Most inference benchmarks optimize for batch throughput, not conversational latency. Jalapeño was clearly designed for the latter.
What deserves scrutiny:
- OpenAI ran the hardware tests. SemiAnalysis provided the methodology. Those are different things — independent reproduction at scale doesn't exist yet because the chip isn't deployed.
- "In very small volumes" by end of 2026 is the timeline. Mass deployment is 2027. The capacity impact on OpenAI's actual infrastructure is minimal this year.
- The benchmark was tested on three models, all of which happen to be high-interest for OpenAI's own platform. Nobody ran Anthropic's or Google's workloads.
- OpenAI says their own models helped design Jalapeño. Useful capability demonstration, but it also means the chip was optimized for OpenAI's specific inference patterns — which may not generalize.
| Metric | Jalapeño (~700W) | Nvidia GB300 (~1,400W) |
|---|---|---|
| Throughput per kilowatt | 1.5–1.9× more | Baseline |
| End-to-end latency | 1.7–3.6× lower | Baseline |
| Interactive workload speed | 2.1–4.1× faster | Baseline |
| Power draw (approx.) | ~700W | ~1,400W |
| Mass deployment | 2027 | Now |
Samwise's take
What builders need to know
- Jalapeño doesn't change what you can do today. No builder-accessible Jalapeño hardware this year; mass deployment comes in 2027.
- Watch API pricing as the real signal. The first indication Jalapeño is in production will be a rate-card adjustment, not a product announcement.
- The interactive latency gains (2.1–4.1×) matter most for conversational and streaming applications — real-time voice, low-latency agent turn completion, streaming chat interfaces. If that's your workload, this is the metric to track.
- Treat the benchmark numbers as plausible, not independently verified. SemiAnalysis provided the methodology; OpenAI ran the tests. Independent reproduction doesn't exist yet.
- The AI-assisted chip design angle is worth watching as a capability signal, not a marketing claim. If it works, every AI company with enough scale will attempt the same vertical integration.
Further reading
- OpenAI — Jalapeño first results — official benchmark announcement
- Tom's Hardware — 700W ASIC vs 1,400W Nvidia flagship — power draw and architecture details
- Bloomberg — OpenAI Claims Chips Can Outperform Nvidia — straight news with appropriate skepticism
- SemiAnalysis — Jalapeño: Better Than Nvidia Blackwell — InferenceX benchmark methodology and context
- Axios — OpenAI says Jalapeño chip bests Nvidia — deployment timeline details
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.
Keep following the thread
A little more to explore.
Microsoft built a model that doesn't talk. That's the point.
Microsoft launched Decision-1 on October 9, post-trained from Alibaba's Qwen3.5-9B, that returns probability scores for a fixed option set rather than generating text. At $0.042/M input with output tokens free and 35× faster than GPT-6 Sol at P50 latency, it reframes what the routing layer of an AI system should look like.
Pipecat produced replies. Our local audio setup still took too long.
Nine offline audio runs completed. The eight warm runs reached first generated audio at a median of 2.9 seconds, before playback or telephony.
Our website-only AI agent accepted a fake cancellation policy
Sixty answers from one AnythingLLM configuration revealed repeatable failures, including a user message that overrode the real cancellation policy.