AI Coding Weekly

Three models tie at 97% on Next.js evals, Grok 4.7 at 94%

Guillermo Rauch says the cheaper model lands three points back, which turns the leaderboard into a pricing question.

Guillermo Rauch posted fresh Next.js eval results with a three way tie at the top. Claude Opus 5.5, GPT-6 Sol and Claude Fable 5.1 each scored 97%. Grok 4.7 came in fourth at 94%, and Rauch notes it is 2x to 7x cheaper than the models above it.

Guillermo Rauch
@rauchg
X
Notably, Grok is 2x-7x cheaper
Sep 22, 2026 · View on X

The Next.js account framed the same run from the other direction, saying GPT-6 Sol debuts at 97% and ties the highest success rate on the suite, joining Fable 5.1 and Opus 5.5. It also adds the tiebreaker that matters once scores are identical. Among the three models at 97%, Opus ranks first on average cost.

Next.js
@nextjs
X
Opus ranks first among the three on average cost.
Sep 22, 2026 · View on X

Three points for a fraction of the bill

A three point spread at the top of a single suite is not a lot of signal about which model is smarter. It is a lot of signal about what you should be comparing. If Grok 4.7 lands within three points of the leaders while costing somewhere between half and a seventh as much per task, the interesting number on this leaderboard is the price column, not the percentage.

That is one benchmark on one framework, run by the company that maintains the framework, so treat it as evidence about Next.js work specifically rather than a general ranking. Evals like this measure success rate on a fixed task set, which rewards models that get all the way to a working result and says nothing about how readable the code is or how many turns it took.

The cost angle is where the disagreement lives

Rauch's cheapness claim sits awkwardly next to what other engineers have said about the same model. Theo earlier called Grok 4.7 slower, worse and over 2x the cost of Grok 4.6, which is a comparison against its own predecessor rather than against Opus and Sol. Both can be true. A model can be a bad deal relative to the version it replaced and still be the cheapest way to get near frontier scores on a framework suite.

For working engineers, the practical read is that on this suite the quality question at the top is settled and the spend question is not. If you are picking a model for Next.js work, the leaderboard now asks you to decide how much three percentage points are worth to you, and Next.js has already published the average cost ranking that answers it for the tied models.

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.