AI Coding Weekly

Zechner flips to Sol the same day Theo says it ran in circles

Two engineers who actually ran the models landed on opposite verdicts in the same week.

Mario Zechner reversed himself on two frontier models inside a day. First he posted that his friendship with Sol had ended and Opus 5.5 was his new best friend. Then he corrected it, putting Sol back on top and saying Opus was having a stroke while thinking.

Mario Zechner
@badlogicgames
X
sol is new best friend, opus is having a stroke while thinking.
Sep 29, 2026 · View on X

Theo got the opposite result on the same pair. Working on ts-rust, his TypeScript compiler port to Rust, he said he threw GPT-6.1 Sol at the job for a few days and it ran in circles and made no progress. Opus 5.5, on the same codebase, decided all of Astra's code was slop and wrote an entirely new crate from scratch, making more progress in 10 hours than Astra had in two weeks. That is the run we covered earlier, and Theo has since corrected his own account of how it happened, since Opus did not pick up where Astra left off.

Theo - t3.gg
@theo
X
I threw GPT-6.1 Sol at this for a few days and it ran in circles and made no progress
Sep 29, 2026 · View on X

Same week, same models, opposite verdicts

Neither of these is a benchmark. One is a compiler port that has been grinding for weeks, the other is whatever Zechner happened to be building when he posted, and he did not say. So the honest read is not that one model is better. It is that two people with a track record of actually running long agent sessions got opposite answers within days of each other, which is what a capability difference by task type looks like from the outside.

Zechner's phrasing is worth noting because it points at where the failure was. Not wrong output, but thinking that goes off the rails, the model burning tokens while looking confused. Theo's complaint about Sol is shaped the same way, circling rather than failing outright. Long-horizon agent work is where these models diverge most visibly, and it is also the hardest thing to evaluate, because the only signal is whether the thing is closer to done after a day.

For working engineers, the practical takeaway is that a single endorsement from someone you trust has a shelf life measured in hours right now, and it is conditioned on their task. Zechner's own timeline is the cleanest evidence of that, since he was the one who changed his mind.

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.