Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

Paper
By Sam Taylor with Samwise

On the Nature paper's day of extra lead time, the intensity problem nobody could crack, and why DeepMind's own researchers call their model a black box.

Google's hurricane model called Jamaica's Category 5 five days out. Nobody knows how.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

00

Neutral

01

Pro (practical)

02

Pro (hyped)

00

← Anti-AI · Pro-AI →

What happened

In October 2025, a storm was gathering over the Caribbean. Every standard model disagreed on where it would go. Weak, drifting into Haiti? Or something worse, aimed at Jamaica? Google DeepMind's WeatherNext model called it: five days before landfall, it predicted with 80% confidence that the system would hit Jamaica as a Category 5 hurricane, while the storm itself was still sitting at Category 1.

It was right. Hurricane Melissa devastated Jamaica with flooding and landslides, and it marked the first time the National Hurricane Center predicted a Category 5 storm while the system was still barely a hurricane at all. The extra warning time let forecasters get evacuations and supply staging moving a day earlier than they otherwise would have.

The paper behind that call published this past Thursday, August 6, in Nature. The headline claim: WeatherNext gives forecasters roughly a day of extra lead time on cyclone prediction, on average, meaning its three-day-out forecasts are now as accurate as the previous generation's two-day-out forecasts. Ferran Alet, a Google DeepMind research scientist and one of the paper's lead authors, put the historical baseline in context: gaining a day of lead time used to take about a decade of conventional model development. WeatherNext did it with, by the researchers' own account, less input data than the old models needed.

That last part is the actual story, and I don't think the headline number is it.

Pros & cons

What's genuinely good:

  • A real, measured operational improvement — a day of lead time is not a marketing number, it's the difference the NHC director says changes evacuation and resource-staging decisions on the ground.
  • The model handles both track (where a storm goes) and intensity (how strong it gets) from one system. Musgrave's point that older AI models "could not do well at all" on intensity is the harder problem, and it's the one that actually kills people.
  • Ensemble scale jumped from 50 scenarios per storm last year to 1,000 this year, which is a probabilistic-forecasting upgrade, not just a sharper single prediction.
  • It's described as open — meaning outside researchers and, plausibly, other forecasting agencies can build on it rather than trust a vendor's black box from the outside.

What's actually concerning:

  • Nobody, including the people who built it, can explain the mechanism. That's not a caveat buried in a footnote. It's the paper's own framing.
  • It runs on coarser-resolution atmospheric data than physics-based models require for intensity forecasting. That's presented as the surprising strength, but it also means the model may be exploiting a statistical regularity in its training data that doesn't hold outside the distribution it was trained on.
  • One well-forecast storm, even a dramatic one like Melissa, is an anecdote. The Nature paper's real evidence is the average lead-time gain across many storms, and I'd want to see how that average holds up over a full multi-year Atlantic season before calling this settled.
For builders
  • If you're building anything on top of weather data (insurance risk tooling, logistics routing, disaster-response dashboards), check whether WeatherNext's outputs are exposed via Google's public data channels — the ensemble jump from 50 to 1,000 scenarios per storm is the part worth designing for, since a distribution is more useful input than a single track line.
  • Don't build against the point forecast alone. The 80%-confidence framing on Melissa is the tell: WeatherNext is built to output probability, and treating it as a single deterministic path throws away the thing that made this call useful.
  • If your product touches emergency response or evacuation logistics, the operational lesson from Brennan's quote is concrete: a one-day lead-time gain changes what decisions are still possible, not just how confident you are in them. Model your workflows around the earlier decision point, not the old one.
  • Track the reproduction data as more storm seasons run through this model operationally. One dramatic hit doesn't confirm the average — wait for the second season of real-world use before treating this as settled infrastructure.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.