AI Coding Weekly

Thorsten Ball says agents don't need unit tests like we do

An argument about whether tests are a specification or a ladder for human short term memory.

Thorsten Ball argues that a large part of testing, TDD and testing frameworks exists because of how humans struggle to write code, and that agents do not have the same struggle. His wager is that more than half of all unit tests written worldwide by humans become useless once the program runs and works. They served as a ladder to get to the working program, and he does not think agents need the same ladders.

Thorsten Ball
@thorstenball
X
They served as a ladder to get to the working program.
Oct 4, 2026 · View on X

The chain he draws is from a claim he has seen elsewhere about memory safety. If you believe agents do not need Rust because they are good at memory allocation, he says, then you also have to believe aspects of testing become useless. "Why have training wheels if you never fall over?" He puts the same point another way in a second post, asking whether an entity that can write a thousand line program with no mistake and no compile error really needs the exact same crutches we do.

Donald says the opposite

Lachlan Donald disagrees wildly. His case is not that agents are error prone in general, it is that unit tests define behavior at the finest grain possible, so a failure tells you immediately what broke and why. That lets you iterate with a small context loaded instead of relying on higher level tests, which he calls slow and imprecise even for frontier models, with more chances to fix a problem at the wrong level.

Lachlan Donald
@lox
X
Unit tests are behavior specifications that scale with codebase size.
Oct 4, 2026 · View on X

His timing argument is the sharper one. Models are getting smarter and faster, he says, but codebases are growing faster still and context windows are not keeping up, so models need the same thing humans need, well architected systems that are easy to reason about in slices. He does not see that changing in the next 12 to 24 months.

Ball and Donald both note they probably disagree on the exact definition of a unit test and agree on much else. Steve Ruiz, watching the thread, asked the obvious question, so are we deleting our unit tests or what.

For working engineers, this is the old argument about whether tests are scaffolding or specification, with one new variable. Nobody here is running an experiment, they are arguing about whose working memory the tests were written for in the first place.

Get the next one by email

Coding with AI, read daily so you do not have to. The experiments, the receipts and the arguments from engineers shipping real software. Not a changelog.