I tested my new Claude skill against an agent that didn't have it — and the control group found the saving my skill missed
The skill was 35% faster and used 12% fewer tokens. The baseline agent, with no instructions at all, found €10.51 the skill walked straight past. Why the no-skill run is a second opinion, not just a scoreboard.
My new Claude skill was 35% faster than an agent working without it, used 12% fewer tokens, and tested more discount codes. It also missed €10.51 that the agent without the skill found in about two minutes, by noticing that the same product was cheaper per kilo in a different pack size.
That is the whole reason to run the baseline.
What I built, and why it needed testing at all
I had spent an afternoon manually hunting savings on a €120 grocery cart — testing whether an "app-only" discount code actually works on the web (it does), checking whether Swiss retailers carried the same brand more cheaply (they don't, one had delisted it entirely), and working out the customs threshold under which a food order crosses the border duty-free. It worked, and it was the kind of thing I would do again, so I turned it into a Claude skill: a folder of instructions plus a reference file of Swiss-specific levers that Claude loads when I show it a checkout page.
The obvious next step is to declare victory. The skill works, I used it, it found savings. Except that proves nothing, because I never watched Claude try the same task without it. I was about to record a comparison in which one side never played.
Anthropic's own guidance on authoring skills is explicit about this: build scenarios that test the gaps, measure performance without the skill first, then iterate against that baseline. It reads like process hygiene. It is actually the only thing standing between you and a permanent, confident, unfalsifiable belief that your skill is good.
The run
One evaluation, a different cart from the same shop, €86. Two fresh agents: one with the skill installed, one with nothing but the task. Seven assertions about what a good answer contains.
Both passed all seven. The skill was meaningfully faster — 8.8 minutes against 13.5 — and got there with fewer tokens, because it already knew which retailers to check and did not have to rediscover the checkout-testing method from scratch. On the metric I had set out to improve, it improved things.
And then the baseline agent, wandering around the product catalogue without my instructions telling it where to go, noticed that a large pack of the same product had a lower unit price than the small ones in the cart. A €10.51 saving on an €86 order. My skill had walked past it, because my instructions did not contain a step for it, because when I did the task by hand in the first place I had not thought of it either.
So the skill inherited my blind spot and then executed it 35% faster. That is what a skill is: a recording of one person's approach, with the gaps included.
David McRaney, in You Are Not So Smart, on the argument from ignorance: "No matter how you feel about the question, you would be incorrect to assume the lack of evidence proves your assumption." My skill reported no further savings. I would have read that as there being none.
The control group is doing a different job than you think
I set up the baseline to answer "is the skill better?" — a scoring question, with a yes or no at the end of it. That was the boring half of the result.
The useful half is that an agent without my instructions is an agent that has not been told what to ignore. Every line I write into a skill is simultaneously a shortcut and a fence. The speed gain and the missed pack size are the same mechanism viewed from two sides. Constraining the search is what makes it fast, and what makes it blind.
So the baseline is not really the control condition in an experiment. It is a second opinion from something with no priors. I folded the format-swap lever into the skill the same evening, and the skill is now better than it could have become through any amount of me re-reading my own instructions, because I cannot review my way out of a gap I do not know is there.
This also changes how often I think I should run one. I had it filed as a launch-time gate — build the skill, prove it beats baseline, ship, done. It is closer to a recurring check, in the way that a routine that silently fails for three weeks needs something outside itself to notice. A skill that has been running for six months has had six months of the world changing underneath its instructions, and it will keep reporting success the whole time.
The parts that generalise
Run the task twice, once naked. Not a mental simulation of how it would have gone — an actual run, with an actual agent, that you read afterwards. It costs one evaluation.
Score on your criteria, then read the losing transcript anyway. My scoring rubric had seven assertions and none of them was "found every available saving," so on the scoreboard the two runs tied. The €10.51 was invisible to my own measurement and visible in thirty seconds of reading.
Treat the gap as material, not as an embarrassment. The point of the exercise is not to certify the skill. It is to harvest whatever the unconstrained version did that the constrained one could not.
And keep the distinction in mind between a skill and an agent — I've written about where that line actually sits — because a skill is procedural knowledge, and procedural knowledge is exactly the kind that goes quietly out of date. The same instinct applies to the tooling decisions around it: I only knew that building my own screen recorder beat paying for one because I built the thing and measured it, rather than reasoning about it from the outside.
None of this is an argument against skills. The skill is better than me at this task now, in both directions — faster than my manual process and, since yesterday, in possession of a lever I never used. It got there because I let something dumber have a turn.
Build the skill. Then let the version without it embarrass you once, on purpose, while it is cheap.
Sources & further reading
External
Agent Skills — Claude Platform Docs
Skill authoring best practices — Claude Platform Docs
Related posts
Claude skills vs Claude agents — the difference that actually matters
My scheduled Claude Code routine failed three weeks running
I replaced my Loom subscription with a single app I built myself
The death of generic AI