Pocock wants /retro to turn repeated tasks into CLIs
Fewer tokens and fewer hallucinations, says Pocock. One reply says the risk is a script that quietly fails.
Matt Pocock is considering adding a recommendation to /retro that turns repeated tasks into custom tooling, specifically a CLI plus a skill. His reasoning is that an agent delegating work to scripts it owns and controls spends fewer tokens and hallucinates less. He credits a conversation with @poteto for the idea.
the power of agents delegating tasks to scripts they own and control
The /retro command is Pocock's own, and it already reads back your recent agent sessions looking for patterns. The proposed addition is the obvious next step. If the retro spots the same manual sequence three times, it should not just tell you about it, it should suggest the agent build itself a command for it.
The counterargument
Asmir replied that the only risk is that the scripts rot. The API or repo layout changes, the script quietly fails, and the agent trusts its output more than it would trust its own reasoning. Pocock said he was confused by what rotting scripts meant, which is where the exchange in the sources ends.
The failure mode Asmir describes is specific and worth holding onto. A broken script is not a problem when it errors out. It is a problem when it returns something plausible and empty, because the agent reads that as a finding rather than as a tool failure. A model asked to search a codebase from scratch will at least notice when nothing matches. A model that shells out to its own stale helper gets a clean exit code and moves on.
For working engineers
This is the build versus regenerate tradeoff moved down a level. Cached tooling is cheaper than reasoning right up until it is silently wrong, and the whole appeal of the agent-owned CLI is that the agent stops thinking about that step. That is the saving and that is the exposure, and they are the same thing.
Nothing here has been shipped or measured. Pocock is considering it, Asmir raised an objection, and neither has run the experiment on a real codebase over the months it would take for a script to actually go stale. If you want to try the pattern yourself, the cheap mitigation is making the tools fail loudly rather than returning empty results, but that is a suggestion the sources do not make and nobody in them has tested.
the api or repo layout changes and the script quietly fails, and the agent trusts its output more than it would trust its own reasoning

