Is the test suite an audit or a target? Is there such thing as a comprehensive test suite for software?
"when an AI agent's instruction is "go make the failing test pass," the test suite stops being an audit and becomes the target. A human team passing 46,000 tests will pick up understanding as a byproduct of the work. An agent loop passing them just proves the tests pass. It's the same green checkmark, but it carries less information than it used to."