Theo has been tracking which models he swears at
A joke metric from @argofowl turned out to be something Theo had already been logging.
@argofowl asked whether agents could be scored by how often you have to stop them and ask what they are doing. Theo replied that he had already been doing it.
can we score agents by how often you have to stop them and ask 'what the fuck are you doing
"Funny enough, I've been tracking which models I swear at for a while now," he wrote, adding that the results probably would not surprise anyone. He attached the tally as an image. We are not going to read a ranking off a screenshot and present it as a number, so if you want the order, take it from his post.
I've been tracking which models I swear at for a while now
Why a swear count is not a joke metric
Frustration per session measures something benchmark scores do not. A model can land a correct diff and still burn twenty minutes of your attention getting there, rewriting files you did not ask about, or wandering off mid task while you sit there deciding whether to interrupt. That cost lands on the human, so it never shows up in a pass rate.
It is also the cost that decides which model stays open in your editor. Theo has spent weeks on more formal comparisons, including a $1,000 router benchmark and a five model audit scored by a judge panel. A private log of when he lost his temper is a cruder instrument than either, and it may track daily use better than both.
The obvious caveat is that this is one person's data about one person's reactions across whatever work he happened to be doing. It is not an eval. Theo did not claim it was, and he flagged the results as unsurprising rather than revealing.
Still, the underlying idea travels. If you are choosing between two agents that both finish the task, the one you interrupt less is the one that is actually cheaper, and nobody currently publishes that number. Anyone can start counting.

