Vol. 1 · Edition 036Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

Controversy
By Sam Taylor with Samwise

On the demyelination transcription error, the Prozac hallucination a GP caught in her own notes, and why the MHRA's medical-device decision is the real fault line.

Your NHS AI scribe might already have the wrong diagnosis in your file

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

02

Pro (practical)

00

Pro (hyped)

01

← Anti-AI · Pro-AI →

What happened

A woman in England had an MRI. It came back clean, no nerve damage. But the consultation summary her hospital produced, the one an AI scribe had typed up while she was in the room, said she had demyelination. That's the kind of nerve damage that can be an early sign of multiple sclerosis. She's a health professional herself, so she knew to question it. The hospital came back and said the correct wording was "null demyelination," meaning none found. One dropped word turned a negative result into a diagnosis.

Healthwatch England, the NHS's statutory patient watchdog, is publishing this case and others as evidence that AI scribes, now used by GPs and hospital doctors across at least 27 different products in England, are making mistakes that go unnoticed by the doctors reviewing them. In a separate case, a scribe mixed up a prescribed drug with a different one that had a similar name. In another, an AI-generated letter left out that a consultant had told the patient to get a migraine prescription refilled through their GP. In all three, it was the patient who caught it. Not the doctor.

What's documented vs what's disputed

Documented:

  • Healthwatch England has collected multiple specific patient-reported cases, including the demyelination mix-up, the drug-name confusion, and the dropped prescription instruction
  • The Medicines and Healthcare products Regulatory Agency has decided not to classify AI scribes as medical devices, meaning there's no England-wide safety oversight regime for them
  • A GP, Dr Shier Ziser Dawood, published a case last year in a medical journal where an AI scribe recorded that she'd told a patient to "continue their Prozac," a drug she never discussed or prescribed
  • Patients in Rotherham separately complained to their local Healthwatch that an AI receptionist couldn't understand strong Yorkshire accents
  • The government's 10-year NHS plan explicitly expects AI scribes to cut administrative burden and free up clinician time

Disputed or unresolved:

  • How often these errors happen system-wide. Healthwatch is working from patient complaints, not an audited error rate
  • Whether the errors persist in the permanent medical record when nobody catches them
  • Whether doctors are actually saving time, given Dr Dawood's own point that reviewing every AI transcript for mistakes cancels out a lot of the promised efficiency gain

Timeline

  • 2025: Dr Dawood publishes her warning in a leading medical journal, calling AI scribes potentially "a double-edged sword" for GPs, citing the fabricated Prozac reference
  • Earlier in 2026: Patients in Rotherham raise accent-recognition complaints about an AI receptionist to their local Healthwatch branch
  • August 31, 2026: Healthwatch England publishes its findings, including the demyelination case, and the Guardian reports that ministers have separately been warned the NHS and individual doctors could be sued over AI scribe mistakes

Source spread

The part none of these sources says outright

Here's what I keep coming back to. Every single error in this report was caught the same way: a patient happened to notice. The demyelination patient is a health professional who knew to question her own scan result. Dr Dawood's Prozac case surfaced because she personally remembered she'd never prescribed it. Nobody has a mechanism that catches the errors when the patient isn't paying close attention, doesn't have relevant training, or just trusts the paperwork. Which means the real error rate here isn't just underreported. It's structurally invisible, because the only correction mechanism running right now is the exact human vigilance these tools were sold to reduce.

That's also why the MHRA's decision matters more than it reads on first pass. Not classifying AI scribes as medical devices sounds like a technicality. But "null demyelination" becoming "demyelination" is precisely a diagnosis-and-treatment-relevant error, the kind medical device regulation exists to catch before it ships, not after a patient happens to catch it herself.

For builders
  • If you build or sell clinical documentation AI in the UK, check now whether your product's output touches diagnosis, medication, or treatment instructions. That's the line the MHRA is currently letting scribes sit on the wrong side of, and regulatory guidance on this is likely to tighten
  • Build a patient-facing correction and reporting flow into the product itself. Right now the correction mechanism is "the patient happens to notice," which isn't a mechanism, it's luck
  • Don't ship a scribe feature that silently drops qualifier words (null, no, not) without a confidence flag. That single failure mode produced the worst documented case here
  • If your pitch deck says "frees up clinician time," get a real number for review-and-correction time before you use it. Dr Dawood's experience suggests that number eats into the savings

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.