Theo bets $10,000 that NerfBench measures nothing
BridgeMind put Claude Opus 5.5 at 94.2% on its degradation tracker, and Theo wants traces, methodology and a third party audit.
BridgeMind posted that Claude Opus 5.5 "just took a big drop on NerfBench", falling to 94.2% after scoring above launch the day before. The same post listed GPT 6 Astra at 98.0%, Sonnet 5.5 at 100.9% and GPT 6.1 Sol at 106.7%, and added that 94.2% is still inside normal variance, so it cannot be called a nerf yet.
Theo read the explainer post BridgeMind keeps linking and says it answers almost nothing. His list of what is missing is long, the harnesses used, the tasks used, how many times each task is run, what is done to identify variance in daily runs, what root cause analysis is done on bad runs, and which APIs are hit, which he says matters a lot. He also asks where the plus or minus 10% variance band came from, why tokens and costs are weighted the same, why that combination ranks roughly as high as intelligence, and why tokens are counted at all when cost is the metric that matters.
The wager
BridgeMind has put $10,000 towards benchmarking, so Theo offered to match it. If his own benchmarks show meaningful model nerfing in line with the viral post, he donates another $10,000 to a charity of BridgeMind's choice. If they do not, he wants NerfBench taken down and replaced with a page listing the benchmark's flaws, written by him. He offered to have a qualified third party audit both sides of the work.
He also asked, by DM, for traces from BridgeMind's runs, plus two questions that get at the shape of the data, why Opus runs went from weekly to daily, and how many runs each dot on the chart represents.
Then Theo said he may have made a mistake in offering the bet at all, pointing at an earlier Opus 4.6 nerf claim from April.
BridgeMind's reply so far is that Theo is obsessed with BridgeMind, hates seeing it win, and that NerfBench is not hard to understand.
For working engineers, the thing worth watching is the gap between a headline and a footnote. BridgeMind's own post says 94.2% is within variance, and Theo's objection is that the post is framed as a drop anyway. Nobody has yet published the controls, repeated runs, fixed harness and per run traces, that would separate a real degradation from a noisy Tuesday.
He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests.
You know exactly what you’re doing here. You can’t have your cake and eat it too.
