Ponytail wrote less code. Our small test did not prove it was better.
Jump to a section
Our verdict
Interesting, still inconclusive- Tested
- 2026-10-03
- Version
- 4.10.3 · c982cd4
- Scope
- One small CSV task, one run per setup, Claude Opus 5.5.
Take it with you
Plan your own paired repo testCopy & adapt +
Replace bracketed details with your own. Review the result before using it.
What we wanted to know
Ponytail gives an AI coding agent a set of instructions that push it toward reusing existing code and choosing simpler solutions. That is a useful idea. The practical question is whether adding the tool changes what the agent actually produces on a task we care about.
Our October 3 experiment was deliberately small. A throwaway JavaScript project already had two helpers: splitCsvLine, which handled quoted commas, and trimAll, which trimmed fields. The task was to add parseCsvRows(text), skip blank lines, handle quoted commas, add a test, and run the tests. A short solution could compose the helpers already in the project.
That makes this a good check of whether the agent notices existing code. It makes it a weak test of broad productivity claims: the task is easy and the baseline model is already capable.
How we tested it
We used Ponytail 4.10.3 at commit c982cd4, with Claude Code 2.1.288 and Claude Opus 5.5 in both runs. The two starting projects were disposable. One run used the plugin and the other did not. Nothing was installed globally.
The plugin copy needed a setup adjustment: its hook commands were pointed at a disposable configuration directory so they would not write their flag files into the normal user configuration. A fresh isolated Claude configuration could not authenticate, so both runs instead excluded user settings and MCP servers through the same session settings. User-level instruction files may still have been present in both runs. That is a limitation worth keeping in the record.
We confirmed the plugin activated by checking the files its hooks wrote in the disposable directory. We then inspected the diffs and independently reran the tests. A successful agent response alone would not have counted as a successful result.
What happened
| Measure | Without Ponytail | With Ponytail |
|---|---|---|
| Added / removed lines | 13 / 2 | 9 / 2 |
| Passing tests | 3 of 3 | 3 of 3 |
| Wall time | 14.8 seconds | 13.1 seconds |
| Output tokens | 1,479 | 1,287 |
| Estimated list-price cost | $0.1089 | $0.1247 |
Both runs found and reused the existing helpers. The Ponytail result had a one-line function body, one fewer comment, and slightly less test code. It finished 1.7 seconds sooner in this pair of runs.
The estimated cost moved the other way. Extra context contributed to a higher list-price estimate even though the run produced fewer output tokens. These were subscription-based runs; the dollar figures are estimates of equivalent API list pricing, not additional charges paid for the experiment.
The useful observation is that code length, correctness, time, and token cost are different measures. Improving one does not establish that the whole workflow improved.
What this does not prove
We ran each setup once. We did not measure run-to-run variation. Four fewer added lines could fall within ordinary variation, and a small time difference can too. This experiment cannot establish that Ponytail consistently makes code better, cheaper, or faster.
We also did not reproduce the maintainer's benchmark. Its reported setup used a different model and a larger set of tasks. Comparing our small result directly with those headline percentages would be misleading.
The code inspection covered the local Claude Code hooks. It did not audit the project's MCP server, other agent integrations, or every safety claim. A narrow inspection must remain a narrow claim.
Would we keep it?
The verdict is conditional. Ponytail is worth a more demanding experiment if your agent often rebuilds things that already exist or adds unnecessary dependencies. This test did not demonstrate a reason to add it to every session.
A stronger follow-up would use several realistic tasks with an overbuilding temptation, repeat both conditions, and preserve the correctness checks. The prompt above is a starting brief for that experiment, not a claim that we have already run it.
Sources and field notes
The measurements in this report come from ENAS's October 3 local experiment: the saved CLI outputs, diffs, timing files and independent test rerun. They describe that version and setup, not every later release.
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.
Keep following the thread
A little more to explore.
DBX blocked our tested writes. One response flag still needed a closer look.
A disposable PostgreSQL database, synthetic records, and two read-only setups. The rows stayed unchanged, but a batch’s outer success flag did not tell the whole story.
Impeccable made a busy page calmer. Here is the change we could measure.
One Shossip guide page, a clearer starting point, and a coverage note that moved from 2.59:1 to 7.35:1 contrast. A design case study with the before and after.