An AI-pilled company capped its SOTA model budget
Gergely Orosz talked to a company that had unlimited AI budgets a year ago and a technical CEO bullish on AI. It now runs a daily cap and uses the priciest models for planning only.
The big disagreement
Theo says using Jev to rank reasoning model outputs is its worst use
Theo argues that anything simple enough for Jev to validate is simple enough that most modern models get it right anyway, and that Jev cannot validate much because it cannot run tools or modify its own context. His original complaint was about using a model that does not reason to pick between outputs of models that do.
Ranking outputs from expensive reasoning models might be the worst case I’ve seen for it thus far.
Receipts
Company with unlimited AI budget now caps its priciest models daily
Gergely Orosz says a company that had unlimited AI budgets a year ago has put a daily limit on its most expensive models. Orosz quotes them as using Fable and Astra for planning only, with cheaper models judged good enough for everything else, and he stresses Sol and Opus remain unlimited there.
Use Fable and Astra only for planning, cheaper models are good enough for everything else.
Model and agent watch
dax says matching DeepSeek inference takes $100M, and claimants are lying
dax says his team has been working for months to reproduce DeepSeek's quality of inference and calls it $100M-budget difficult, needing the right connections and the right ten people in the world. It follows his claim that providers advertising 99% cache rates are wrapping DeepSeek, since he says DeepSeek is the only provider hitting that number.
if you ever see anyone claiming they can do this, they are lying


