AI Coding Weekly

Theo's Terminal Bench 4 run puts Sol past Opus 5.5 on cost

Same model, same benchmark, wildly different numbers depending on whether you run Codex or mini-swe-agent.

Theo says he ran Terminal Bench 4 on GPT-6.1 Sol himself because there were no benchmarks yet when he made his video, and the scores came back better than Opus 5.5 for roughly a thirtieth of the price.

Theo - t3.gg
@theo
X
Performance better than Opus 5.5 for ~1/30th of the price
Sep 29, 2026 · View on X

Artificial Analysis then published its own run. Theo's summary of it is that Sol performs at around Opus 5.5 Medium levels for under a third of the price, lower than his numbers but with what he calls similarly insane costs, which puts Sol on the pareto frontier for the task either way.

The harness is the variable

The reason for the gap, as Theo reads it, is scaffolding rather than the model. He used Codex. Artificial Analysis used mini-swe-agent, which he describes as "think pi but shit". In a follow-up run he compared against numbers from Harbor Hub, which ran each model on its official harness, and found every model scored higher there than in the AA run. Sol gained the most of any model from moving off mini-swe-agent.

That is one engineer's runs, not a controlled study, and Theo himself says the difference is big enough that he might redo Terminal Bench 4 on the official harnesses.

What he actually uses it for

The cost headline has not changed his defaults for writing code. Theo still calls Opus 5.5 the current GOAT for coding and says he does not reach for Sol there. He frames Sol instead as a replacement for Astra at a much better price, and says he does default to it for code reviews, architecture analysis, computer use, email management and a lot of his other daily work.

For working engineers, the useful takeaway is not the ranking, it is that a single Terminal Bench figure describes a model plus a harness. Two organisations ran the same benchmark on the same model in the same week and landed on different conclusions about how close it gets to Opus 5.5, because one wrapped it in an agent loop the model does well in and the other did not. If you are picking a model off a leaderboard, check which scaffold produced the number, and whether it looks anything like the agent you actually ship.

Theo - t3.gg
@theo
X
All models see a bump compared to AA, but 6.1 Sol sees the biggest bump moving from mini-swe to codex
Sep 29, 2026 · View on X

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.