Theo says smart and dumb are two axes, not one
A model can bench as brilliant and still do things that make no sense, and Theo says that is not a contradiction.
Theo argues that we should stop treating "smart" and "dumb" as opposite ends of one axis when talking about language models. We do it because that is how we think about people. Models, he says, do not work that way. A model can be incredibly smart and incredibly dumb at the same time, and the two properties vary independently.
Astra is still the smartest model available today, but it is also one of the dumbest, regularly doing things that make literally no sense whatsoever.
His first example is Gemini. Theo describes those models as incredibly smart, with a level of baked in knowledge he calls incredible and capability you can see directly in benchmark results. Then you hand one a job. The amount of stupid things the models will do, he writes, is similarly incredible. High capability, high stupidity, same model, same session.
His second example is the comparison most people arguing about model choice actually care about. Theo puts Fable 5.1 as not quite as smart as Astra, but significantly less dumb. Astra, in his framing, sits at the top of the capability axis and near the top of the other one too.
Why this reframing is useful
If smartness and stupidity are one axis, then a leaderboard ordering is a shipping recommendation, and the model at the top should be the one you put in your agent loop. Engineers keep finding it is not. The ranking captures the capability axis reasonably well. It does not capture how often the model does something that makes no sense on a real codebase, which is the thing that costs you an afternoon of review.
This is the reliability versus capability tradeoff restated in a form that is easier to act on. Under the one axis model, a smart model doing something dumb reads as a fluke or a bad prompt. Under two axes, it is just the second coordinate, and you can pick for it. Theo frames the practical question as finding the right balance of both, and says it is a challenge for users and researchers alike.
Worth being clear about what this is. It is a framing argument from one person, not a measurement, and the placements of Gemini, Fable 5.1 and Astra on his two axes are his read rather than a scored result. Nobody in the sources is offering a stupidity benchmark. But it matches the shape of what engineers keep reporting when a model that tops the charts turns out to be the one they trust least unattended, and Theo's claim is that once you stop insisting on either-or, understanding model behavior gets much easier.
