ThePrimeagen swapped Luna for Haiku 5.5 and it got worse
A one line config change was the only difference, which rules out the harness as an explanation.
ThePrimeagen tested Claude Haiku 5.5 against GPT-6 Luna on some automation work, both with reasoning set to none, and says Haiku performed significantly worse. It was also significantly slower, he says, on both passes and failures.
Haiku, literal drop in replacement (it was a one line change of changing a string in a config file), performed significantly worse
The detail that makes it worth reading is how little changed between the two runs. He swapped one string in a config file, moving from openai/gpt-6-luna to anthropic/claude-haiku-5.5, with reasoning left at none for both. That is his own stated reason for posting the config, so nobody could write the result off as an artifact of how he was calling the model.
He has not explained the gap. "Still have to understand why", he wrote, which is the honest position given he ran one workload and has one set of results.
Against the pitch
Anthropic announced Haiku 5.5 as "the cheapest, fastest, and most capable small model we've ever released" and said it costs around 75% less to run on average than Claude Haiku 4.5. That framing is a comparison against the previous Haiku, not against anything from another lab, and price per token is only half the arithmetic. A cheaper model that needs more attempts, or takes longer per attempt, can cost more in wall clock time and in retries than the headline number suggests.
That is the shape of ThePrimeagen's complaint. He did not report a pricing comparison at all. He reported that the drop in replacement was slower whether it succeeded or failed, which is the worst case for an automation loop, because the failures cost you time before you even find out they are failures.
For working engineers, the useful part is the method rather than the verdict. Small fast models get picked for exactly this kind of job, mechanical automation where you do not want to pay for reasoning, and the usual way people evaluate them is by reading the launch post. Holding the harness and the reasoning setting fixed and changing a single identifier is the cheapest experiment that tells you whether a swap is actually free. One engineer, one workload, one result, and he says he does not yet know the cause. Worth running on your own pipeline before you trust the 75% figure to mean 75% less spend.
On average, it costs around 75% less to run than Claude Haiku 4.5.

