AI Coding Weekly

Theo's judge panel ranks Astra first, Grok 4.7 far back

Then he ran the same ranking with Astra as judge, and Fable finished last.

Theo gave the same big audit task to all the new models, then set up a judge panel to rank the resulting audits. Astra came out on top, with Opus and Fable close behind. Grok 4.7 and GPT-6 Sol, in his words, were WAY behind. He also posted a cost breakdown alongside the rankings, noting he pays for subscriptions to all of them.

Theo - t3.gg
@theo
X
Astra slaughtered. Opus and Fable close behind.
Sep 24, 2026 · View on X

The catch is in his own follow-up. Those rankings were produced by a Fable panel. When Theo had Astra generate a similar report over the same audits, Astra put Fable at the bottom.

The judge is part of the result

That single swap is the whole story. The ordering did not change because the audits changed, it changed because the grader changed. A model asked to rank prose it did not write still brings its own idea of what a good audit looks like, how long it should be, how it should be structured, how confident it should sound. Swap the judge and you swap the taste.

This is not an argument that the exercise is worthless. Theo ran the experiment, which puts him ahead of anyone theorizing about it, and the gap he reports between the top three and the bottom two is wide enough that it survived at least one judge. It is an argument for reading the result as one person's setup with one grader rather than as a benchmark number you can quote back at someone.

What is actually being measured

An audit is close to the worst case for automated scoring. There is no test suite that passes or fails, no diff that applies cleanly. Whether a security or code audit is good depends on whether the findings are real and whether the important ones are in there, and checking that requires someone who knows the codebase. A judge panel is measuring something adjacent to that, and how adjacent is unknown.

Separately, Theo said Astra is still the best review model and called it insanely thorough, which is his impression from use rather than an output of this ranking.

The Grok 4.7 placement lines up with what Theo said earlier about the model. That is consistency in one person's testing, not independent confirmation.

If you want to copy the method, copy the second half of it too. Run the ranking with at least two different judges and see whether the order holds. If it does not, you have learned something about your judges, which is worth knowing before you wire one into a pipeline.

Theo - t3.gg
@theo
X
Astra is still the best "review" model. It's so insanely thorough.
Sep 24, 2026 · View on X

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.