Inference optimization
Techniques and costs for matching model quality, speed, and caching performance. Every AI Coding Weekly story on inference optimization, newest first. 3 stories since September 2026.
- Bun AoT starts Claude Code 24% faster, says Jarred Sumner
The experimental build also cuts the Claude Code install from 430 MB to 304 MB, with a nearly 2x larger binary.
- Anthropic says it made claude.ai 3x faster in two weeks
The write-up includes the prompts and methods, which is the part worth testing on your own codebase.
Have the good opinion, not the loud one
One email a day. Under three minutes. What working engineers found out about coding with AI.
- dax says matching DeepSeek inference is a $100M problem
He also says any provider advertising a 99% cache rate is wrapping DeepSeek, because DeepSeek is the only one hitting that number.
People on Inference optimization
Get the next one by email
Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.