Theo's bench: Sol beats Opus 5.5 at 1/30th the cost
Theo ran Terminal Bench 4 himself and got GPT-6.1 Sol above Opus 5.5 for roughly a thirtieth of the price, then spent the same day telling people he still would not use it to write code.
Receipts
Zechner switches to Sol the same day Theo says it ran in circles
Mario Zechner flipped positions inside a day, first posting that Opus 5.5 was his new best friend and then correcting himself in favour of Sol, saying Opus was having a stroke while thinking.
sol is new best friend, opus is having a stroke while thinking.
Model and agent watch
Theo's own Terminal Bench 4 run puts Sol above Opus 5.5 on cost
Theo says he ran Terminal Bench 4 on GPT-6.1 Sol himself because no benchmarks existed yet, and got performance better than Opus 5.5 for about a thirtieth of the price. He credits the gap with Artificial Analysis's lower scores to the harness: he used Codex, AA used mini-swe-agent, and Sol gained the most of any model from the switch.
Performance better than Opus 5.5 for ~1/30th of the price
Gergely Orosz: OpenAI locks in companies by refusing to lock them in
Orosz argues OpenAI is pulling ahead of Anthropic on enterprise strategy because Codex is open source, runs other models, and can be used with your own harness, none of which Claude Code allows.
To lock in mid-sized and above companies, you need to NOT lock them in...


