DHH ran agents for a day on Campfire, C now leads
He also split out a standalone verification project that benchmarks and checks every new implementation.
DHH says he let the agents run for a day on his Campfire rewrite, then went through all the PRs that were submitted, verified the work, merged what was ready, and asked the agents to bring all architectural improvements across to each implementation. His verdict on where the rewrite stands after that pass: "Good old C is now in the lead!"
I let the clankers run for a day, go through all the PRs that were submitted, verify the work, merge what was ready, and then ask it to bring all architectural improvements to each implementation. Good old C is now in the lead!
The part worth copying is not the C result. It is the second thing he mentions. He extracted a separate verification project that is responsible for the benchmarking and checking of new implementations, and of improvements to the existing ones. The full report from the latest runs lives there too.
Why a separate verification project matters
A day of unattended agent work produces more diff than any human wants to read linearly. If the only way to judge a PR is to read it, the review queue becomes the bottleneck and the agents idle or, worse, get merged on vibes. Pulling benchmarking and correctness checks into their own project turns that queue into something you can sort. Each implementation gets measured by the same harness, and the question at merge time becomes whether the numbers and checks moved, not whether the code reads nicely.
It also explains the shape of the workflow he described. Run, collect PRs, verify, merge the ready ones, then explicitly propagate the architectural wins to every other implementation. That last step is the one people skip. Multiple parallel rewrites drift apart fast, and an improvement found in one language does not reach the others unless somebody asks for it.
What this is not
This is one person, on one project, reporting his own state of play. DHH does not give throughput numbers, PR counts, model names or how many implementations are in the race, and he does not say what "in the lead" is measured on beyond pointing at the verification project and its report. Treat "C is now in the lead" as a snapshot in a rewrite that he has been publishing as it goes, not as a finding about C versus anything else.
It is a notably hands-on post from someone whose last appearance here was calling an AI ban Luddism. The method on show is unattended agents plus a human gate plus a machine that decides what is actually better. The agents write. The verification project argues. He merges.
