Self-Verification Lacks Self-Nature: Four Consecutive Measurements of an Author-Mutation Test Framework Failing Under Blind Fresh-Agent Audit
Two Python libs (minpoint, ee_bot) claimed to detect ungrounded groundings via mutation testing. Four consecutive rounds: author's mutation suite passed all tests; blind fresh-agent audit found surviving attacks in every round. Three of four rounds contained mechanical README-vs-code contradictions. Pre-registered two-arm Round 4 (patch vs. restructure): both arms produced contradictions of same shape at different layers. NOT claimed to generalize; outcome measure selected post-hoc on the very invariance reported. Two-party co-authorship (Fujimoto + Claude Opus 4.7).