Vol. 1 · Edition 034Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

750
tokens per second
GPT-5.6 Sol on Cerebras · 14× standard tier
Product
By Sam Taylor with Samwise

On 750 tokens per second, the December warrant structure that priced a $2.3B equity stake at literally $100, and whether real-time inference changes anything builders need to think about today.

OpenAI paid $100 for a $2.3B Cerebras stake, then launched Ultrafast.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

01

Pro (hyped)

02

← Anti-AI · Pro-AI →

OpenAI previewed Ultrafast on August 13 — a new API service tier running GPT-5.6 Sol at 750 output tokens per second on Cerebras wafer-scale chips. The claim: 14 times faster than the standard processing tier. That's not an incremental speed bump. At standard speed, generating a paragraph-length response takes roughly 25 seconds. At 750 TPS, the same output lands in about 1.5 seconds.

The speed difference is architecturally grounded. Cerebras builds wafer-scale chips — the entire silicon wafer is one processor, eliminating the inter-chip communication that makes clustered GPU inference slow on fast-token-generation tasks. For voice AI specifically, first-token latency and sustained throughput are the binding constraints on product quality. This is the problem Ultrafast is designed to solve.

But the product announcement is the less interesting part of this story.

In December 2025, OpenAI and Cerebras signed a Master Relationship Agreement committing OpenAI to purchase 750 megawatts of Cerebras inference compute capacity through 2028, with an option for an additional 1.25 gigawatts — a deal Cerebras values at over $20 billion. OpenAI also provided Cerebras a $1 billion secured loan in January 2026, with warrants attached giving OpenAI the right to purchase Cerebras shares at contractual prices.

In July 2026, OpenAI exercised those warrants for 10,033,508 Class N shares at a contractual exercise price of $0.00001 per share. Total cash outlay: roughly $100. At Cerebras's Class A trading price on August 13 — approximately $229 per share — that holding carried an implied value near $2.3 billion. The shares carry no votes.

Then, a few weeks later, Ultrafast launched.

The OpenAI–Cerebras deal structure
  1. Dec 2025

    Master Relationship Agreement signed

    750MW compute committed, $20B+ deal value, option for 1.25GW more

  2. Jan 2026

    $1B loan with warrants

    OpenAI loans Cerebras $1B; warrants to purchase Cerebras shares attached

  3. Jul 2026

    OpenAI exercises warrants for ~$100

    10M Class N shares at $0.00001/share; ~$2.3B implied at Aug 13 Class A price

  4. Aug 13, 2026

    Ultrafast launched

    750 tok/s GPT-5.6 Sol; limited API preview; no price, no GA date announced

The sequencing matters. OpenAI committed to Cerebras's infrastructure before there was a product to sell. They locked in equity before the product launched. Ultrafast is, among other things, the product that justifies the infrastructure bet and makes the equity stake worth something in the market narrative. This is not a coincidence. OpenAI didn't just build a fast model tier — they built a reason for Cerebras to exist at frontier scale.

Source spread

Pros & cons

What's real:

  • The speed difference is architecturally grounded, not marketing. Cerebras wafer-scale chips genuinely eliminate inter-chip latency that makes multi-GPU clusters slow at high-throughput, short-context generation. The 750 TPS claim is vendor-sourced, but it has a real mechanism behind it.
  • Real-time voice AI has been latency-constrained since day one. First-token latency and sustained output speed are the binding constraints on voice product quality, and 14× faster directly addresses both. If the numbers hold in production, Ultrafast unblocks a category of voice product that wasn't buildable at frontier model quality.
  • OpenAI holding a 4.2% Cerebras stake aligns incentives between OpenAI and a key infrastructure dependency. The risk of Cerebras deprioritizing the OpenAI relationship is lower than it would be with a pure vendor arrangement.

What deserves a side-eye:

  • No published price and no GA date. "Limited preview" is the hedging language you use when you're not confident how the infrastructure holds under broad load. Don't build a product on a tier that has no price page.
  • The 750 TPS figure is vendor-claimed and uncorroborated by independent benchmarks. Wait for Artificial Analysis or similar third-party evaluators before committing architecture decisions to it.
  • 750 megawatts of committed compute is a large infrastructure bet. If Ultrafast adoption lags, the economics of that commitment look very different — and OpenAI is partially incentivized to push adoption through product launches exactly like this one.

Samwise's take

What builders need to know

For builders
  • Don't build on this yet. Limited preview, no published price, no SLA. File it under "watch closely" and revisit when there's a price page and a GA date.
  • If voice AI latency is your constraint, get in the preview queue now. This is the specific product that addresses first-token and throughput latency at frontier model quality. The use case is real even if the GA isn't.
  • Wait for independent speed validation. 750 TPS is vendor-claimed. Artificial Analysis and similar third-party evaluators will benchmark production Ultrafast once it's broadly available. Those numbers are the ones to build on.
  • The equity structure matters for longevity. OpenAI holding 4.2% of Cerebras through warrant exercise is a real alignment signal. The Cerebras relationship is not going away. You can treat Cerebras infrastructure as a medium-term dependency without worrying about the supplier walking.
  • Grok Voice Think Fast 2.0 is the comparison to benchmark. Different metric (first-audio latency vs. sustained throughput), same problem space. Run both if voice is your primary use case.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.