Kenton Varda says testers mistook Opus 5.5 for a new Fable
A reviewer who says he still reads the code rates Opus 5.5 close to Fable on style, and keeps GPT models for bug hunting.
Kenton Varda says a group of early access testers were convinced Opus 5.5 was a new Fable, and were shocked when told otherwise. He posted the impression after a weekend of coding with it, calling the model's clearer speech its best feature, with far less invented jargon.
In early access a bunch of us were convinced it was a new Fable, totally shocked when they told us it was Opus.
my loop for a while has been Fable codes, Sol/Astra reviews
The framing matters more than the verdict. Varda says he cares a lot about the quality of code written by AI because he still reads it, which puts his judgment in a different category from a pass rate on a test suite. He is also explicit that he has not rigorously compared the two models. Everything Opus 5.5 produced last weekend seemed pretty good to him, and he rates it similar in quality to Fable but with clearer comments.
The GPT complaint is about taste, and he says so
Varda has not been happy with the code quality of any GPT model, including Astra. His objection is not correctness. The code works, he says, he just does not like the way it looks, and he hates the total lack of comments. He allows that this may simply be a matter of taste.
Fable was the first model where he felt happy with almost all the code it output, usually needing only to edit some comments for clarity. That is the bar Opus 5.5 is being measured against here, a bar about readability rather than whether the tests go green.
The split loop
The transferable part is the workflow. Varda has frequently found GPT models better at spotting and diagnosing bugs, so his loop for a while has been Fable writing code and Sol or Astra reviewing it. Different families for different jobs, chosen by what each one is good at rather than by which one currently tops a leaderboard.
He has not done much bug finding with Opus 5.5 yet, so he cannot say whether it is catching up on that axis. That gap is worth holding onto. A model that writes code you enjoy reading is not automatically the model you want interrogating a subtle failure, and Varda is careful not to claim otherwise.
For working engineers, this is one person's impression from one weekend, stated as such by the person who had it. The useful takeaway is not the ranking but the question behind it. If you still read what the agent writes, style and comment quality are real acceptance criteria, and they are not what the benchmarks measure.
