AI Coding Weekly

Theo says Grok 4.7 is slower, worse and over 2x the cost of 4.6

The promised token efficiency gain went the other way, and Theo says he cannot find a single benchmark where it holds.

Theo ran Grok 4.7 on real engineering tasks and called the release disappointing. His summary is that it is slower, less pleasant to use, scores worse than Grok 4.6 on various benchmarks, and costs more than 2x Grok 4.6 in real-world use.

Theo - t3.gg
@theo
X
The frontend capabilities are unacceptably bad. The 3D capabilities are nonexistent.
Sep 22, 2026 · View on X

The sharpest part is the token efficiency claim, because it was the headline promise. Elon Musk said in advance that Grok 4.7 would be the 2.1T model, "better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency." Theo says he has yet to find a single bench where 4.7 is more token-efficient than 4.6, and puts the regression at 30 to 80%.

Theo - t3.gg
@theo
X
I have yet to find a single bench where Grok 4.7 is more token-efficient than Grok 4.6.
Sep 22, 2026 · View on X

The cost math

Token efficiency is the number that decides the bill. A model priced per token can look cheap and still cost more per task if it burns more tokens getting there, which is what Theo says is happening here. He also notes that on Artificial Analysis, Grok 4.7 comes out more expensive than Astra.

That matters more than another benchmark row. Theo's account of the line is that Grok 4.5 was fast, reliable and a solid default model for the price, and that 4.6 was a forgivable step in the wrong direction, slower and more expensive, using way more tokens per task for a slight edge in intelligence. A model earns default-model status on price and speed, not on the top of a leaderboard, and if 4.7 has moved above Astra on cost then the reason anyone reached for Grok in the first place is gone.

The parts benchmarks miss

Theo is explicit that the scores are not the whole story, and says 4.7 is pleasant to use in various real-world engineering tasks while still feeling "so 2025". His specific complaints are frontend and 3D work, which he calls unacceptably bad and nonexistent respectively, plus getting stuck in random loops.

This is one person's runs, not a controlled evaluation, and it is worth holding it at that weight. But the cost claim is the checkable one, and it is the one that changes behaviour. Anyone who set Grok as the cheap default in an agent loop can measure tokens per completed task on their own workload before and after, and find out in an afternoon whether Theo's 30 to 80% shows up for them too.

Theo ends by hoping the SpaceXAI team acknowledges the miss and impresses with the next release.

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.