AI Coding Weekly

Theo vs NerfBench, and Opus 5.5's widening lead

Theo offered to match BridgeMind's $10,000 to settle whether NerfBench measures anything, and BridgeMind refused to share the transcripts. Meanwhile Theo's own July-to-September scorecard flipped hard toward Anthropic.

The big disagreement

BridgeMind refuses Theo's $10,000 bet and will not share NerfBench traces

The NerfBench fight ended with no bet: BridgeMind declined to hand over run transcripts, promised more methodology later, and told Theo to go build his own bench. Theo had offered to match BridgeMind's $10,000, with a third party audit and a public apology page written by whoever lost, after listing nine things the methodology post never says, including harnesses, tasks and how many runs each point represents.

BridgeMind
@bridgemindai
X
It's about your ego, and not wanting anyone else to exist in this space.
Oct 3, 2026 · View on X

Read the full story

Receipts

Theo flips: preferring OpenAI models for code now makes almost no sense

Theo closed the week with a July versus September scorecard and says preferring OpenAI models for code now makes almost no sense, after months of arguing the opposite. His vibe ratings put Opus 5.5 at 9/10 on code capability and 9/10 on understanding intent, against 7.5 and 3.5 for GPT-6 Astra.

Theo - t3.gg
@theo
X
It made more progress in 10 hours than Astra did in 2 weeks
Sep 28, 2026 · View on X

Read the full story

How they actually work

Pocock promotes /retro from prompt of the day to default step

Matt Pocock moved /retro into the main flow of his skills, so session retrospectives are now a standing step rather than a tip he posted once. The prompt reads his last 10 coding agent sessions and looks for places where agents took too long to find relevant information or relied on out of date docs.

Read the full story

Model and agent watch

T3 Code ships cross-provider delegation and native subagents in one PR

Theo landed a T3 Code release whose headline feature is `delegate_task`, letting an agent start child agents on any provider or model. It also adds an ACP registry for agents like Devin, Cline, Kimi and Droid, native subagents shown as child threads with their model, status and history, mid-thread provider switching, thread forking and auto-resume when limits reset.

Read the full story

Open tabs

DHH: every developer needs an always-on AI shed

DHH says every developer needs an AI shed, an always-on box on their tailscale network running most of their agents.

Vercel confirms a KVM zero-day through its sandbox bounty

Guillermo Rauch says Vercel confirmed a KVM 0day through its Sandbox bounty program, in what he calls the industry's gold standard for Linux virtualization. A full writeup is coming.

Frequently asked questions

What is NerfBench and why does it matter?

NerfBench is a methodology for detecting model nerfing (making a model perform worse). It matters because nobody has published a reproducible method for it yet, so claims that a model got worse are still just assertions without proof.

Which coding model should I use according to recent benchmarks?

Theo recently changed his recommendation from OpenAI models to preferring Opus 5.5 for code writing, rating it 9/10 on code capability and 9/10 on understanding intent. He recommends GPT-6.1 Sol as a cheap replacement for code reviews and architecture analysis.

What is the /retro prompt and how is it used?

The /retro prompt reads your last 10 coding agent sessions and looks for places where agents took too long to find relevant information or relied on outdated documentation. It's now a standing step in the workflow rather than an optional tip.

What new features did T3 Code add for multi-agent work?

T3 Code added delegate_task which lets an agent start child agents on any provider or model, native subagents shown as child threads, mid-thread provider switching, thread forking, and auto-resume when limits reset. These features let you build trees of agents spread across different providers.

What is an AI shed?

An AI shed is an always-on box on a developer's tailscale network (a private network) that runs most of their agents locally.

Built from 4,180 posts by the engineers, language designers and agent builders we follow on X over 7 days.

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.