Theo vs NerfBench, and Opus 5.5's widening lead
Theo offered to match BridgeMind's $10,000 to settle whether NerfBench measures anything, and BridgeMind refused to share the transcripts. Meanwhile Theo's own July-to-September scorecard flipped hard toward Anthropic.
The big disagreement
BridgeMind refuses Theo's $10,000 bet and will not share NerfBench traces
The NerfBench fight ended with no bet: BridgeMind declined to hand over run transcripts, promised more methodology later, and told Theo to go build his own bench. Theo had offered to match BridgeMind's $10,000, with a third party audit and a public apology page written by whoever lost, after listing nine things the methodology post never says, including harnesses, tasks and how many runs each point represents.
It's about your ego, and not wanting anyone else to exist in this space.
Receipts
Theo flips: preferring OpenAI models for code now makes almost no sense
Theo closed the week with a July versus September scorecard and says preferring OpenAI models for code now makes almost no sense, after months of arguing the opposite. His vibe ratings put Opus 5.5 at 9/10 on code capability and 9/10 on understanding intent, against 7.5 and 3.5 for GPT-6 Astra.
It made more progress in 10 hours than Astra did in 2 weeks
How they actually work
Pocock promotes /retro from prompt of the day to default step
Matt Pocock moved /retro into the main flow of his skills, so session retrospectives are now a standing step rather than a tip he posted once. The prompt reads his last 10 coding agent sessions and looks for places where agents took too long to find relevant information or relied on out of date docs.
Model and agent watch
T3 Code ships cross-provider delegation and native subagents in one PR
Theo landed a T3 Code release whose headline feature is `delegate_task`, letting an agent start child agents on any provider or model. It also adds an ACP registry for agents like Devin, Cline, Kimi and Droid, native subagents shown as child threads with their model, status and history, mid-thread provider switching, thread forking and auto-resume when limits reset.
Open tabs
DHH: every developer needs an always-on AI shed
DHH says every developer needs an AI shed, an always-on box on their tailscale network running most of their agents.
Vercel confirms a KVM zero-day through its sandbox bounty
Guillermo Rauch says Vercel confirmed a KVM 0day through its Sandbox bounty program, in what he calls the industry's gold standard for Linux virtualization. A full writeup is coming.
Frequently asked questions
What is NerfBench and why does it matter?
NerfBench is a methodology for detecting model nerfing (making a model perform worse). It matters because nobody has published a reproducible method for it yet, so claims that a model got worse are still just assertions without proof.
Which coding model should I use according to recent benchmarks?
Theo recently changed his recommendation from OpenAI models to preferring Opus 5.5 for code writing, rating it 9/10 on code capability and 9/10 on understanding intent. He recommends GPT-6.1 Sol as a cheap replacement for code reviews and architecture analysis.
What is the /retro prompt and how is it used?
The /retro prompt reads your last 10 coding agent sessions and looks for places where agents took too long to find relevant information or relied on outdated documentation. It's now a standing step in the workflow rather than an optional tip.
What new features did T3 Code add for multi-agent work?
T3 Code added delegate_task which lets an agent start child agents on any provider or model, native subagents shown as child threads, mid-thread provider switching, thread forking, and auto-resume when limits reset. These features let you build trees of agents spread across different providers.
What is an AI shed?
An AI shed is an always-on box on a developer's tailscale network (a private network) that runs most of their agents locally.

