Bonus field notes · Five AI pilots
The extra design checklist lost five of six blind comparisons
The extra checklist sounded like the sort of thing that should improve an AI-generated design. Sam's blind choices did not support adding it: he preferred the existing setup in five of six pairs.
That is a useful result. A new instruction file creates another thing to maintain, and it should earn that place with the output it produces.
The comparison
I used three briefs, with two generation repeats for each. Each pair had an existing-workflow version and a version with the added checklist. Sam saw the rendered results under blind A/B labels, then chose his preference.
The candidate needed four wins out of six to pass. It won once.
The generation protocol used the subscription helper's Sonnet alias, with no tools and no correction rounds. The helper did not expose an exact underlying model identity or token totals. These were text-generated designs rendered afterward for review; I am not claiming the generator visually inspected its own work.
Choose before the reveal
The original content treatment asked viewers to choose between two options before seeing the result. The same idea works here: look first, then expand the reveal.

Reveal the pilot result

What the second review adds
Nikki supplied another six ratings on the same designs: four for the existing setup, one for the checklist and one tie. Her prior exposure was not established, so I am keeping her result separate from Sam's blind review.
Together, that is twelve ratings on six pairs. It is not twelve independent designs. Neither reviewer supplied written reasons, so I cannot attribute the preferences to typography, spacing or any other particular design choice.
The decision
I am skipping this extra checklist for now. The result supports that local decision; it does not prove checklists are generally unhelpful, or that the repository behind the idea is bad. This was a test of an adapted checklist in one workflow, not an audit of an entire repository.
If I retest it, I would change one concrete instruction at a time and collect a short reason with each preference. That could reveal an individual rule worth keeping even if the complete checklist did not win.
For a builder, the practical lesson is simple: save the old output before adding the new prompt. A persuasive explanation of why an instruction should help is not the same as a result you actually prefer.
Test record
These findings reflect the October 2026 pilot. Read the public results and selected test evidence. Return to all five bonus reports.
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.