On the AA-Briefcase Elo jump from 1,313 to 1,577, what 53 turns vs 103 means for token budgets at scale, and whether this belongs in your agent stack at $2/M.
SpaceXAI held the model size constant and cut agentic turn count in half. That's the bet.
Grok 4.6, released August 12, runs the same 1.5 trillion-parameter V9 base as Grok 4.5. SpaceXAI spent the intervening month on post-training — regenerated SFT data, extended RL in agentic environments — and got 5 Artificial Analysis Intelligence Index points and a roughly 2× reduction in average turns per long-horizon task. The turn-efficiency result is underweighted in the coverage.
Read the full take →