Vol. 1 · Edition 039Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

GPT-6 AstraFable 5.1Grok 4.7 (xAI)Grok 4.7 (AA)
Model Launch
By Sam Taylor with Samwise

On the 2.1T new base model, Terminal-Bench 4.0 at 26–38%, and what a 5× price gap to Fable 5.1 and GPT-6 Astra actually buys you.

Grok 4.7 is bigger, cheaper, and behind the frontier. That's xAI's whole argument.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

02

Neutral

00

Pro (practical)

02

Pro (hyped)

00

← Anti-AI · Pro-AI →

SpaceXAI released Grok 4.7 on September 21 with a new base model: 2.1 trillion parameters, up 40% from Grok 4.6's 1.5T V9 architecture. The training run was longer, weighted toward tasks that take hours to complete, and produced what xAI calls self-verification improvements — the model double-checks its own outputs more often than 4.6 did. Price stayed flat: $2 per million input tokens, $6 per million output. Same as Grok 4.6. Same as Grok 4.5 before it.

Then the benchmark numbers arrived.

Artificial Analysis measured Grok 4.7 at 26% on Terminal-Bench 4.0 — the updated version of the agentic coding benchmark that most builders use for model selection. xAI's own self-reported figure is 38% on the same benchmark. Either way, the comparison is the same: GPT-6 Astra sits at 60% on Terminal-Bench 4.0, Claude Fable 5.1 at 55%. The gap to the frontier is 17 to 34 points depending on which number you start with.

The 12-point discrepancy between xAI's and Artificial Analysis's measurements deserves its own note. Different harnesses, reasoning settings, tool configurations, and token budgets move these scores around. That's real — it's happened with other models too. But a 12-point spread on a score in the 26–38 range means the evaluation conditions matter substantially for Grok 4.7 specifically, more than I'd expect from a well-calibrated model. I'd treat both figures as bounds until independent reproductions settle it.

Terminal-Bench 4.0 vs. API Pricing (September 2026)
ModelTB 4.0 ScoreInput / 1M tokensOutput / 1M tokens
GPT-6 Astra60%$10.00$50.00
Claude Fable 5.155%$10.00$50.00
Grok 4.7 (xAI self-report)38%$2.00$6.00
Grok 4.7 (Artificial Analysis)26%$2.00$6.00

Source spread

Pros & cons

What's real:

  • A new 2.1T foundation model is genuinely different from a post-training release. Grok 4.6 was SFT and RL on the same V9 base. Grok 4.7 is a new base entirely. That matters for how the improvement compounds — and for whether the xAI "longer RL on harder tasks" thesis gets a real test.
  • At $2/$6/M, Grok 4.7 is 5× cheaper than Fable 5.1 and GPT-6 Astra on input, and the same 8× cheaper on output. For workloads where this capability tier is sufficient, that price gap is real money at scale.
  • Availability in GitHub Copilot, Cursor, Grok Build, and the xAI API on day one is a distribution improvement over earlier Grok releases.
  • Self-verification is genuinely hard to capture in a single benchmark score. A model that double-checks its own work more often is quietly better at long-horizon tasks in ways that may not show up cleanly in Terminal-Bench 4.0 solo runs.

What deserves a side-eye:

  • A 12-point gap between xAI's self-reported 38% and Artificial Analysis's measured 26% — on the same benchmark, the same model — is a flag. Labs reporting better numbers than independent evaluators isn't new, but this spread warrants watching until more reproducible results arrive.
  • 26–38% versus GPT-6 Astra's 60% is a tier difference, not a gap that price alone resolves. This isn't a model that's close to the frontier.
  • xAI has been making the "longer RL runs on harder tasks will close the gap" argument since Grok 4.5. The argument is coherent. The gap is still there. The thesis may be right — but it hasn't been proven across three consecutive releases.

Samwise's take

What builders need to know

  • TB 4.0 ≠ TB 2.1. Grok 4.5 scored 83.3% on Terminal-Bench 2.1 back in September. Grok 4.7's 26–38% on Terminal-Bench 4.0 is a harder benchmark. The two numbers aren't comparable; don't let press coverage imply the model got dramatically worse.
  • Run your own eval before making a tier decision. A 12-point spread between xAI's and AA's measurements signals this model is sensitive to evaluation setup. Set your own harness against your actual tasks before committing.
  • If frontier-level output is required, this isn't that. Fable 5.1 and GPT-6 Astra are measurably ahead on Terminal-Bench 4.0. Buy according to what your tasks actually need.
  • If you're cost-constrained, evaluate seriously at this price. $2/$6/M at Grok 4.7 capability is a real offer, especially for high-volume workloads where task-pass-rate differences of 20 points translate to a small fraction of overall output and cost matters.
  • GitHub Copilot access is real. Grok 4.7 rolled out to Copilot on September 21. That's a legitimate way to test it in your coding workflow without any API integration overhead.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.