OpenClaw deleted 400k lines of its own tests
Peter Steinberger says OpenClaw removed around 400k lines of its own test code with little change in coverage, and the trick was giving the agent a number to hit instead of telling it to clean up.
The big disagreement
Opus 5.5 edits files with bash, and Armin and Mario disagree on fighting it
Armin Ronacher says his current experience with Opus 5.5 is the model reaching for bash to edit files instead of the edit tools, and relays Mario Zechner's joke about writing an extension that tells it to stop.
don't fight the model, give in to its RL, let if flow over you
Receipts
Steinberger: OpenClaw deleted 400k lines of tests, coverage held
Peter Steinberger says OpenClaw deleted around 400k lines of its own tests without much change in code coverage, blaming models that write tests for every tiny change. The technique is the prompt: instead of asking the agent to clean up, he gives it a target like removing 20% of the least useful tests while keeping coverage within 2%.
If you just tell the agent to clean up, it will stop far too early.
Theo ran five models through one audit and judged them with a panel
Theo had the new models each produce a big audit, then set up a judge panel to rank the audits: Astra won, Opus and Fable close behind, Grok 4.7 and GPT-6 Sol well back. The rankings came from a Fable panel, and when Theo had Astra rank the same audits, Astra put Fable last.
Astra slaughtered. Opus and Fable close behind.


