dax says agents keep accidentally undoing functionality
Nothing in a repo records that you rejected solutions one through four and picked five.
dax says the slop problem is mostly behind us, and the next one is people accidentally undoing functionality. His diagnosis is social before it is technical. Everyone is now in everyone else's business, so when someone hits an issue they fix it immediately instead of nagging the person who wrote it, and nothing in the repo tells them how intentional the current behavior was or what absolutely should not be rolled back.
where is the source of truth for how your software is supposed to work?
The part that makes it specific is the case tests cannot cover. You tried five approaches, four of them looked valid, and you shipped the fifth for reasons that live in your head or in a thread somewhere. A test asserts that the behavior exists. It does not assert why approach three was rejected. As dax puts it, someone later, and he means an agent, might think "oh 3 is simpler" and quietly take you back to a solution you already threw out.
He also pre-empts the obvious fix. Writing another document does not feel like a solution to him, because the document is one more artifact that drifts away from the code. His question at the end of the post is the honest version of the problem. Where is the source of truth for how your software is supposed to work.
Pocock says this is a solved genre
Matt Pocock replied that this is what ADRs are for, meaning architecture decision records, short dated documents that log a decision and the reasoning behind it. His framing is the sharper line in the thread. Code is the materialized view of all the decisions made, so you have to keep track of the decisions or the agent will unmake them.
Code is the materialized view of all the decisions made
That is an answer to the question but not obviously an answer to dax's objection, since an ADR is also a document that can drift. Nobody in the exchange shows a setup that keeps the record current without a human maintaining it, and neither of them reports running an experiment here. This is two practitioners arguing about where intent should live, not a result.
For working engineers, the thing that changed is throughput, not the problem. Undocumented intent has always decayed, and it was survivable when the people capable of rolling it back were humans who moved slowly, asked around and had some context about what the surrounding code was for. An agent that reads the repo and nothing else has no way to distinguish a load bearing choice from an accident, and it will refactor both with equal confidence. If you have ever answered a code review comment with "we tried that, it breaks on Windows", that sentence currently exists nowhere a model can find it.

