Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

AnthropicOpenAIGDMMetaZ.aiAlibabaxAIDeepSeekMistral
Safety
By Sam Taylor with Samwise

On nine labs graded across six domains, the pause pledges four of them quietly voided, and what it means to build on infrastructure when the industry keeps moving its own goalposts.

The AI industry's safety report card is out. The class valedictorian got a C+.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

03

Pro (practical)

00

Pro (hyped)

00

← Anti-AI · Pro-AI →

A C+ is not a good grade. It falls between "adequate" and "slightly better than adequate" in most systems. It is not the grade you tell people about.

Anthropic earned it.

The Future of Life Institute released its Summer 2026 AI Safety Index earlier this month — nine major AI labs evaluated across six domains by an independent panel of seven experts from the University of Montreal, UC Berkeley, and Oxford University. Anthropic received the highest grade. At C+.

OpenAI and Google DeepMind both received a C. Meta landed at D+. Z.ai and Alibaba Cloud received a D−. And xAI, DeepSeek, and Mistral all failed.

That last part is worth sitting with. Three of the most-deployed AI providers globally — including Elon Musk's xAI and two of China's dominant open-weight houses — received failing grades from one of the field's most credible independent evaluators. Not "needs improvement." Not "room to grow." Failing.

What the domains actually measured

Six domains: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and transparency and communication.

Anthropic led in five of the six. OpenAI led in Risk Assessment — "on the strength of a broader evaluation suite and diverse engagement with external testing," per the panel. That distinction is real. OpenAI's pre-deployment red-teaming has gotten more systematic, and the panel recognized it.

The weak link, industry-wide, is existential safety. No lab earned above a C− on that domain. Most scored D or below. The panel's assessment of whether the industry has credible plans for what to do if its most capable models turn out to be genuinely dangerous: poor, across the board.

FLI AI Safety Index Summer 2026: Nine Labs
LabGradePause Pledge Status
AnthropicC+Weakened — contingent on competitors
OpenAICWeakened — contingent on competitors
Google DeepMindCWeakened — contingent on competitors
MetaD+Weakened — contingent on competitors
Z.aiD−No prior pledge on record
Alibaba CloudD−No prior pledge on record
xAIFNo prior pledge on record
DeepSeekFNo prior pledge on record
MistralFNo prior pledge on record

The finding nobody leads with

Grades aside, the one I can't stop returning to: four of the highest-ranked labs — Anthropic, OpenAI, Google DeepMind, and Meta — have all weakened or voided prior commitments to pause AI development unilaterally if their systems approached specified danger thresholds.

The panel assessed that this pattern had undermined safety frameworks across the board. Their framing: these labs have moved the goalposts on their own stated red lines.

The reasoning from the labs, where it's been articulated, is roughly: "We can't stop unilaterally if competitors won't." Which is understandable as a business constraint and alarming as a safety posture. A commitment that only holds if everyone keeps it simultaneously is not a commitment. It's a coordination problem wearing a commitment's clothes.

The military use expansion is the slower-moving story. Between 2024 and 2026, labs that had previously prohibited offensive military applications successively reversed those policies and expanded defense cooperation. This happened without a policy crisis. It is now the industry norm.

Source spread

Pros & cons

What's real:

  • An independent panel — not industry self-assessment — evaluated nine labs across six distinct domains. That's more rigorous than most AI safety coverage, which runs on press releases.
  • Anthropic's C+ reflects genuinely stronger performance on transparency, governance, and safety frameworks versus everyone else graded. The lead is meaningful.
  • OpenAI leading in Risk Assessment is a specific, verifiable advantage if pre-deployment evaluation rigor is what you're optimizing for in an API provider. The panel recognized something real there.
  • The existential safety finding is a factual observation: no lab has credible documented plans for what to do if its most capable model turns out to be genuinely dangerous. That gap is real.

What deserves a side-eye:

  • Letter grades carry false precision on something as contested as "AI safety." Two panels with different methodologies would produce different grades. The ranking is more reliable than the specific letter.
  • "Anthropic leads" is a comparative claim. Comparative claims about safety are dangerous when the comparison group is mostly failing or getting Ds.
  • There's no enforcement mechanism. A company can receive an F and continue operating. The index is informational, not regulatory, and labs know it.
  • The gap between C+ and C is the smallest delta on the chart. The meaningful cliffs are between D− and F, and between C and D+. The top four labs are clustered tighter than the grades suggest.

What builders need to know

  • The pause pledge erosion is a real risk signal for infrastructure planning. The labs you're building on have decided their most important safety commitment only holds if competitors agree. That tells you something about how they'll resolve future tradeoffs between safety and competitive pressure.
  • Existential safety is the binding constraint. No lab scored above C− on the domain that asks "what happens if this goes badly wrong." The gap between "deployed today" and "credible plan for serious failure modes" is real across the entire industry.
  • xAI, DeepSeek, and Mistral received failing grades. If you're deploying these in high-risk environments — healthcare, financial services, government — the index should be part of your vendor due diligence documentation. Not because F means "don't use," but because your risk register should reflect it.
  • OpenAI's Risk Assessment lead is actionable. Broader evaluation suites and more diverse external testing are verifiable differentiators, not marketing claims. That matters if you're choosing APIs for regulated deployments.
  • Military use policy shifts will filter into commercial access. Labs that have expanded defense partnerships may eventually face model classification or access restrictions on their most capable outputs. Start watching model card updates and terms of service for signals.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.