Vol. 1 · Edition 036Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

90.8%
Terminal-Bench 2.1
Gemini 3.8 Flash · Sep 2, 2026
Model Launch
By Sam Taylor with Samwise

On the 90.8% Terminal-Bench score that tops Fable 5.1, Google's three-releases-in-seven-weeks pace, and why January 1, 2027 is the date worth marking.

Google's AI just cleared Claude on the coding test. The pricing window is still running.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

01

← Anti-AI · Pro-AI →

If you use Google products that have AI built in (Gemini in Gmail, the AI overview at the top of Search, Google Docs), the underlying model got an upgrade on September 2. You probably didn't notice. The responses just got a little sharper.

Gemini 3.8 Flash launched September 2, the third major Flash model Google has shipped since July 21. The headline claim: 90.8% on Terminal-Bench 2.1, a benchmark for how well a model handles structured developer coding tasks. For comparison, Claude Fable 5.1 scored 85.02% on the same test last week. Grok 4.5 hit 83.3% on an earlier Terminal-Bench version. That 5-point gap over the current leading model is real, assuming the benchmark holds up to scrutiny.

The pricing is $0.75 per million input tokens and $3.75 per million output, intro rates that double January 1, 2027. Four months from now.

90.8%
Gemini 3.8 Flash on Terminal-Bench 2.1 — above Fable 5.1's 85.02% and Grok 4.5's 83.3% on the same class of test

→ Source: Google blog

Source spread

Pros & cons

What's real:

  • 90.8% Terminal-Bench 2.1 is the highest published number for this benchmark version. The test is designed to approximate real developer tasks and hasn't been saturated at this point.
  • Three major Flash releases since July 21. 3.6, 3.7, 3.8. Each has been genuinely different from the last. Version 3.7 was already the fastest model on Artificial Analysis's tracker at 340.1 tokens per second. These aren't marketing bumps.
  • $0.75/$3.75 intro pricing is competitive for this benchmark tier. Builders can get coding performance above Fable 5.1 at a lower per-token price.
  • Google's retroactive price cuts on 3.6 when 3.7 launched suggests willingness to compete on economics alongside capability.

What deserves a side-eye:

  • 90.8% is a first-party benchmark number from a launch post. I haven't seen independent reproduction of Gemini 3.8 Flash on Terminal-Bench 2.1. The Fable 5.1 comparison is also first-party. Both numbers use the same benchmark, which makes the comparison clean — but neither has been independently verified on this version.
  • "Intro pricing" ends. January 1, 2027 is closer than it sounds. Systems built around $0.75/$3.75 should model what happens when that doubles.
  • Terminal-Bench 2.1 tests structured coding tasks. Production agentic work involves long sessions, ambiguous instructions, and partial failure recovery. Leading the benchmark does not automatically transfer to leading in production.

Samwise's take

What builders need to know

For builders
  • Eval now, not in December. $0.75/$3.75 intro pricing doubles January 1, 2027. You have four months to decide whether Flash belongs in production at the current economics.
  • 90.8% Terminal-Bench 2.1 is the highest published number for this benchmark version, clearing Fable 5.1's 85.02%. Worth revisiting any workflow where you defaulted to Fable 5.1 on coding quality.
  • Reproduce before flipping production. First-party benchmark numbers are a starting point. Run the tasks that actually matter to your use case before committing.
  • Check for behavior drift. Three rapid Flash releases mean prompts tuned for 3.7 behavior may need review for 3.8. Test before assuming compatibility.
  • Verify the knowledge cutoff. Gemini 3.7 Flash moved it to March 2026; the 3.8 announcement doesn't re-state it. Check the Gemini API docs before depending on recent knowledge.

What to do about it (for everyone else)

Google's AI tools in Search, Gmail, and the Gemini app have been getting noticeably better this year, and this upgrade is part of why. The improvements happen automatically. Nothing to install or configure.

One thing worth keeping in mind regardless of which AI product you use: smarter models still make things up. This is called hallucination (the model generates a confident-sounding answer that happens to be wrong). Gemini 3.8 Flash is better than its predecessors, and it still hallucinates. For anything consequential — health questions, financial decisions, legal matters — treat the AI answer as a starting point and verify it with a real source before acting on it.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.