
Theo - t3.gg
Every AI Coding Weekly story that cites Theo - t3.gg (@theo), newest first. 25 stories since September 2026.
On X: Full time CEO @t3dotcodes & @t3dotchat. Part time YouTuber, investor, and developer
- ThePrimeagen swapped Luna for Haiku 5.5 and it got worse
A one line config change was the only difference, which rules out the harness as an explanation.
- Theo ships tsc-rs, a Rust TypeScript compiler he never read
Five months of agent work, two token bills that differ by a factor of twenty, and zero lines of human review.
- Theo fixes a GitHub annoyance with a 90 second Chrome extension
One prompt, one small browser extension, and a point about what fixing other people's software now costs.
- T3 Code adds cross-provider delegation and native subagents
One pull request brings child threads, mid-thread model switching and an ACP registry for Devin, Cline, Kimi and Droid.
- BridgeMind turns down Theo's $10,000 NerfBench bet
The fight over whether anyone can measure model nerfing ends with no transcripts and no bet.
Have the good opinion, not the loud one
One email a day. Under three minutes. What working engineers found out about coding with AI.
- Theo flips, says preferring OpenAI models for code barely adds up
Two months of vibe ratings, and the one that moved most was not capability.
- Theo. Opus 5.5 takes half of all T3 Code prompts
The first model in Theo's coding product to cross a 50% share of traffic, by his own account.
- Theo hides running agent threads in T3 Code
A beta option surfaces an agent thread only when it wants attention, and Theo says the cost is the spatial stability he spent effort building.
- Zechner flips to Sol the same day Theo says it ran in circles
Two engineers who actually ran the models landed on opposite verdicts in the same week.
- Theo says a $200 Claude sub bought $9,000 of Opus usage
His own accounting of his own accounts, not a number from Anthropic.
- Theo says Opus 5.5 threw out Astra's Rust port, not fixed it
The 10 hour result stands, but the mechanism behind it was a from-scratch rewrite in a new crate.
- Opus 5.5 beat Fable 5.1 on HAProxy port time and cost
Two models did the same C to Rust port, and the gap showed up in hours and dollars rather than pass rates.
- Theo's Jev router benchmark shows no win on cost or speed
A $1,000 overnight run on DeepSWE put OpenRouter's new router level with GPT-6 Astra on low, slightly pricier and almost 5x slower.
- Theo does the math on Opus 5.5 usage limits
Half of a Claude Code subscription was off limits to Fable, and Opus 5.5 on high costs over 2x less.
- Theo says smart and dumb are two axes, not one
A model can bench as brilliant and still do things that make no sense, and Theo says that is not a contradiction.
- Theo has been tracking which models he swears at
A joke metric from @argofowl turned out to be something Theo had already been logging.
- Theo's judge panel ranks Astra first, Grok 4.7 far back
Then he ran the same ranking with Astra as judge, and Fable finished last.
- Theo says no lab has both small and large models worth using
A three line post ranks three labs by what they are missing, and shows no work.
- Theo says Grok 4.7 is slower, worse and over 2x the cost of 4.6
The promised token efficiency gain went the other way, and Theo says he cannot find a single benchmark where it holds.
- Theo lists six problems with Jev compaction, Armin Ronacher pushes back
A demo that scores every tool call and deletes the rest turned into the week's argument about what a fast classifier can actually decide.
- Theo says ranking reasoning outputs is Jev's worst use
Malte Ubl counters with the Lean proof argument, that checking work is easier than doing it.
- Claude Code now falls back to AGENTS.md
Thariq says the fallback is live and toggleable, and the replies moved straight on to skills.
- roon says Astra's code makes corrigibility a matter of faith
Theo quote-posted it as OpenAI employees inventing words for unreadable code, then agreed the fix belongs in the model.
- dax moves part of his team off Astra after spend doubles
One team's cost numbers, a few bad impressions and a paused top tier all landed in the same week.
- Mitchell Hashimoto's whiteboard defense for AI code
He does not care who typed the code, only whether you can defend the system in a hallway conversation.
Topics Theo - t3.gg comes up in
Get the next one by email
Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.