Theo bets $10,000 against NerfBench
BridgeMind says Opus 5.5 dropped to 94.2% on its NerfBench tracker. Theo says he cannot tell what the benchmark measures, and is offering $10,000 to settle it.
The big disagreement
Theo offers $10,000 to settle whether NerfBench measures anything
BridgeMind posted that Claude Opus 5.5 fell to 94.2% on its NerfBench tracker, against 98.0% for GPT 6 Astra and 106.7% for GPT 6.1 Sol, while calling that inside normal variance.
He ran 6 of 30 tests in a benchmark, and when 2 failed he claimed "NERF" because he failed to run the other 24 tests.
Receipts
Half of all T3 Code prompts now go to Opus 5.5
Theo says Opus 5.5 is the first model to break 50% of traffic in T3 Code, with literally half of all prompts going to it. No other model has crossed that line in his product.
How they actually work
Theo hides running agent threads and says 18 at once stopped feeling crowded
Theo shipped a beta T3 Code option that hides threads while they are working, so a task only appears when it wants attention. He says it felt weird at first, then he ran 18 threads without it feeling claustrophobic. The tradeoff he flags is losing the spatial stability he had deliberately built in, since things now move around.
Things appear when they need your attention, and disappear when they don't.
