AI Coding Weekly

Theo says no lab has both small and large models worth using

A three line post ranks three labs by what they are missing, and shows no work.

Theo posted a three line verdict on the current model field. Anthropic has no small models worth using. OpenAI has no large models worth using. Google has no models worth using.

Theo - t3.gg
@theo
X
Google has no models that are worth using right now.
Sep 23, 2026 · View on X

That is the entire post. No benchmark, no task, no codebase, no side by side. He did not say what counts as small or large, and he did not name a single model in any of the three sentences. The third line is the blunt one.

What is actually being claimed

Read structurally, the post is less a ranking of labs than a claim about coverage. Two of the three labs get credit for one half of the size range and a fail on the other, which implies Theo thinks the small model tier and the large model tier are different products with different buyers, and that no single lab currently serves both. The Google line breaks the pattern by refusing to split at all.

The size split is the interesting part and also the part with the least support behind it. Small models are the ones you put in a loop, in a classifier, in a hot path where latency and price decide whether the feature ships at all. Large models are the ones you hand a refactor to and walk away. A lab can be excellent at one and unremarkable at the other, and if that is true it changes who you buy from per task rather than per vendor.

How much weight to put on it

This is a taste call from someone who tests a lot of models in public, not a result. Nobody in our window has published numbers that support or contradict any of the three lines, and Theo did not publish any himself. Treat it as a prompt to check your own small model tier against a competitor's rather than as a finding.

If someone runs the comparison and shows the work, that is the version of this story worth acting on. Nobody has yet.

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.