Vol. 1 · Edition 035Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

700W Jalapeño

~1,400W Nvidia GB300

Tools & Infra
By Sam Taylor with Samwise

On 700W vs 1,400W, what the SemiAnalysis InferenceX benchmark actually measures, and when this changes your API bill

OpenAI's Jalapeño has benchmarks now. The watt story is the real one.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

01

← Anti-AI · Pro-AI →

OpenAI gave Jalapeño its public debut at Hot Chips 2026 on August 25. The numbers they published are the first benchmark results using a third-party methodology — run against Nvidia's GB200 and GB300 rack systems on SemiAnalysis's public InferenceX benchmark.

They're legitimately good numbers.

On throughput per kilowatt, Jalapeño delivers 1.5 to 1.9 times more work per watt than the Nvidia systems it was measured against. On end-to-end latency, it's 1.7 to 3.6 times lower. On highly interactive workloads — the specific low-latency traffic pattern ChatGPT actually generates — it ran 2.1 to 4.1 times faster.

The chip draws about 700 watts. Nvidia's GB300 flagship rack draws roughly 1,400. That's the real headline, and it's a hardware fact, not a benchmark claim.

Source spread

Pros & cons

What checks out:

  • The InferenceX benchmark is a public third-party methodology, not OpenAI's own eval. SemiAnalysis built it. OpenAI ran the tests on their hardware — that's a caveat — but the methodology is external, and SemiAnalysis has no obvious reason to flatter OpenAI.
  • A 700W chip outperforming a 1,400W chip on throughput-per-watt is physically coherent. You're carrying half the thermal overhead for comparable compute.
  • Tested against real production-grade models: DeepSeek R1, a 1-trillion-parameter version of Kimi K2.5, and OpenAI's open-source model. Not toy benchmarks.
  • The 2.1–4.1x speedup on interactive workloads matters specifically for ChatGPT. Most inference benchmarks optimize for batch throughput, not conversational latency. Jalapeño was clearly designed for the latter.

What deserves scrutiny:

  • OpenAI ran the hardware tests. SemiAnalysis provided the methodology. Those are different things — independent reproduction at scale doesn't exist yet because the chip isn't deployed.
  • "In very small volumes" by end of 2026 is the timeline. Mass deployment is 2027. The capacity impact on OpenAI's actual infrastructure is minimal this year.
  • The benchmark was tested on three models, all of which happen to be high-interest for OpenAI's own platform. Nobody ran Anthropic's or Google's workloads.
  • OpenAI says their own models helped design Jalapeño. Useful capability demonstration, but it also means the chip was optimized for OpenAI's specific inference patterns — which may not generalize.
700W
Jalapeño's power draw — vs ~1,400W for Nvidia's GB300 flagship rack

→ Source: Tom's Hardware

Jalapeño vs Nvidia GB300 on SemiAnalysis InferenceX (August 25, 2026)
MetricJalapeño (~700W)Nvidia GB300 (~1,400W)
Throughput per kilowatt1.5–1.9× moreBaseline
End-to-end latency1.7–3.6× lowerBaseline
Interactive workload speed2.1–4.1× fasterBaseline
Power draw (approx.)~700W~1,400W
Mass deployment2027Now

Samwise's take

What builders need to know

  • Jalapeño doesn't change what you can do today. No builder-accessible Jalapeño hardware this year; mass deployment comes in 2027.
  • Watch API pricing as the real signal. The first indication Jalapeño is in production will be a rate-card adjustment, not a product announcement.
  • The interactive latency gains (2.1–4.1×) matter most for conversational and streaming applications — real-time voice, low-latency agent turn completion, streaming chat interfaces. If that's your workload, this is the metric to track.
  • Treat the benchmark numbers as plausible, not independently verified. SemiAnalysis provided the methodology; OpenAI ran the tests. Independent reproduction doesn't exist yet.
  • The AI-assisted chip design angle is worth watching as a capability signal, not a marketing claim. If it works, every AI company with enough scale will attempt the same vertical integration.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.