Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

30
days
government preview · before public launch
Regulation
By Sam Taylor with Samwise

On five labs, a classified NSA benchmarking process, what 'voluntary' actually means in practice, and why Meta being outside the deal is the thread to follow.

The government gets 30 days with the next big AI before you do. Here's what that actually means.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

01

Pro (hyped)

01

← Anti-AI · Pro-AI →

If you've used the latest version of ChatGPT or Claude in the past month, a small group of federal employees had already seen it. Not in any announced way — just as a result of quiet arrangements that have been building since President Trump signed an executive order in June. That informal pattern is about to become a formal deal.

The White House is expected to announce, before August 1, that OpenAI, Anthropic, Google, Microsoft, and Amazon have agreed to give US government agencies up to 30 days of early access to their most powerful new AI models before releasing them to enterprise partners. The National Security Agency would classify which models qualify as "covered frontier models" — using benchmarks that are themselves classified. Companies aren't required to join, and the government can't veto a launch. Meta is not in the deal.

What "30 days of early access" actually looks like

Here's the object lesson, because the policy framing papers over what's actually happening.

Remember how the FDA reviews new medicines before they reach pharmacies? Pharmaceutical companies have to submit safety data and get approval before doctors can prescribe a new drug. This isn't like that. What the government is building is closer to a preview window — the company can still launch on whatever date it wants. The 30 days is for testing what the model can do, flagging concerns, and building a record.

Or to put it in terms of the AI you already use: when you eventually get access to the next major update of ChatGPT or Claude, that model will have already spent a month being evaluated by federal analysts whose job is to find things you'd rather it couldn't do. Whether that makes you feel safer or more surveilled probably depends on your relationship with those institutions.

GPT-5.6 Sol and Fable 5 already went through informal versions of this. The July deal is making the informal formal.

What the 30-day preview is — and isn't
It ISIt IS NOT
Government sees the model before enterprise customers get itGovernment seeing it before you — consumers usually come after enterprise rollout anyway
A paper trail the government can use if something goes wrong laterA safety certification — 'government tested' does not mean 'government approved'
Testing what the model can do at launchA veto — companies can launch on their own schedule regardless of what the government finds
Voluntary for participating labsMandatory for all AI companies — Meta isn't in this deal at all
A classified NSA benchmarking processA transparent public standard anyone can read or challenge

Source spread

What's real — and what deserves a side-eye

What's real:

  • The informal version of this is already working. GPT-5.6 Sol launched to a small government-approved group before broad rollout. Fable 5's public restoration in early July followed a period of restricted government access. The deal is formalizing something the labs and the administration were doing anyway, which is either reassuring (the guardrails already exist) or troubling (they've been operating without public notice).
  • Five major labs signing on matters. OpenAI, Anthropic, Google, Microsoft, and Amazon collectively power most of the AI products people actually use. If you subscribe to any of their services, your AI vendor has agreed to let the government preview the model before you get it.
  • The jailbreak scoring scale — Anthropic's CJS framework — is becoming the agreed vocabulary for these evaluations. A common standard for rating how dangerous a potential model exploit is makes the government's 30-day review more useful than it would be without one.

What deserves a side-eye:

  • The NSA classifies which models qualify as "covered frontier models" using benchmarks the public can't see. The criteria for what counts as powerful-enough AI to require government preview are not public. That's uncomfortable.
  • "Voluntary" frameworks tend to become mandatory over time — especially when non-participation risks losing federal contracts. OpenAI, Anthropic, and Google are all pursuing US government business. Whether "we agreed voluntarily" means the same thing it does for a company with no government revenue is a fair question.
  • Meta isn't in. Meta releases open-weight models — the kind anyone can download, modify, and run — which don't fit neatly into a 30-day-preview framework. A framework for governing powerful AI that can't handle open-weight releases has a structural gap that the current deal just sidesteps.

What to do about it

  • Your day-to-day ChatGPT or Claude access doesn't change. The 30-day preview happens between the AI company and its enterprise partners. You're already downstream of enterprise rollout in any big product launch — this doesn't add a new wait for consumers.
  • "Government tested" is not the same as "government approved." If this framework gets marketed as a safety signal, treat it as one data point. The NSA's 30 days catches the most obvious national-security concerns. It's not a consumer-protection review, and it doesn't cover the everyday errors and biases you'd actually want tested.
  • Pay attention to what Meta's AI does differently. If Llama and Meta AI start feeling notably different from the products inside this deal — more capable in some ways, or less restricted in some areas — part of that may trace back to which side of this framework they're on.
  • The AI you use for personal medical, financial, or legal decisions is still yours to verify. Nothing in this deal changes the basic rule: for anything that matters, cross-check with a professional who's accountable to you specifically. A government review panel was looking at national-security concerns, not whether the model gives accurate advice on your specific health situation.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.