Prove nothing changed
September 2026
Michael Feathers defined legacy code in Working Effectively with Legacy Code (2004) as, in one reader's summary, "simply code without tests". The same answer quotes the book's reasoning: "Code without tests is bad code. It doesn't matter how well written it is; it doesn't matter how pretty or object-oriented or well-encapsulated it is. With tests, we can change the behavior of our code quickly and verifiably. Without them, we really don't know if our code is getting better or worse."
Twenty-two years later an agent can open a thousand refactor PRs overnight. Each comes with a tidy description and a green lint run. None of them can answer the one question that matters on a legacy codebase: did behavior change?
My claim is that verification, not model capability, is what unlocks agents on legacy code, and that the usual order is wrong: point a swarm at an old codebase, then review what comes back. On code without tests, review is guessing at scale. The first thing to build is a digital twin: the real backend and the real app running end to end on machines you control, with every user journey written down as something a machine can replay. Once that exists, a swarm can generate millions of versions of the code, and the end-to-end suite decides which ones are correct.
Legacy code is its own specification
Feathers's later essay on characterization testing has the line I keep coming back to: "When a system goes into production, in a way, it becomes its own specification. We need to know when we are changing existing behavior regardless of whether we think it's right or not."
His procedure is almost a joke. Write a test named x, assert that the function returns null, run it, and paste whatever it actually returned into the assertion. "The purpose of characterization testing is to document your system's actual behavior, not check for the behavior you wish your system had." The Stack Exchange answer names what this gives you: the previous version serves as a "consistency oracle since it compares the consistency between two versions." It checks that the code is the same, and says nothing about whether it's correct.
That's the right frame for agents. The agent doesn't know what the code is supposed to do, and often neither do we. What we can know is what it does now.
Old code also carries a lot of weight that does nothing. Romano, Vendome, Scanniello and Poshyvanyk found that "on average dead methods are 15.94% of methods of an application, while the median is 11.95%" across 35 Java GUI apps, with dead methods in 32 of the 35. They also cite an industrial .NET system where 25% of method histories were dead, and an industrial PHP subsystem where developers removed 2,740 dead files, about 30% of it.
Dead code is the easy case for deletion, and deletion is a capability (Simplicity does not sell makes that argument). But you can't delete any of it, or ship the latency work you care about, without near-certainty that the critical flows still work. On an app that moves money, "probably fine" isn't a release criterion. Agents can write the changes. Nothing checks them.
Build the twin before the swarm
"Digital twin" comes from manufacturing. In their 2017 chapter, Michael Grieves and John Vickers trace the idea to a slide Grieves showed at the University of Michigan on December 3, 2002, titled "Conceptual Ideal for PLM." It had a real space, a virtual space, and links carrying data from the real one to the virtual one and information back. It was taught as the Mirrored Spaces Model, became the Information Mirroring Model in Grieves's 2006 book, and the name "Digital Twin" was attached later "by reference to the co-author's way of describing this model." The co-author was Vickers, at NASA. NASA's 2012 paper by Glaessgen and Stargel defined it as "an integrated multiphysics, multiscale, probabilistic simulation of an as-built vehicle or system that uses the best available physical models, sensor updates, fleet history, etc., to mirror the life of its corresponding flying twin."
The chapter has an idea I like more than the name. Grieves and Vickers call it front running: a simulation that runs ahead of the real system and shows operators where their current action leads. They note that "the front running capability requires computing capability that can run faster than the physical activity it is simulating." That's the whole requirement for software too. A twin is useful in proportion to how much faster than production it can tell you what will happen.
The software version is less grand. It's the real backend and the real mobile app running end to end locally, with seeded data, driven by a tool like Maestro, which describes itself as UI automation "using intuitive YAML flows." The tooling matters less than the map. Model the app as a state machine over its screens, and each user journey becomes a path through that graph. The journeys that gate every PR are a chosen subset of those paths.
Once the app is a graph, coverage stops being a feeling. You can ask which transitions no journey crosses, and which journeys carry the most real traffic. I'd weight by usage, because the most-used journeys (on a payments app, paying and signing up) are where a regression costs money.
Evaluation then becomes a function: a device and a journey go in, and metrics, videos and logs come out. A failed submission goes back to the agent with its logs, so it can try again.
The goal I have in mind is a loop that generates millions of versions of an app, each trying to make the code smaller or faster, with every one checked end to end on every journey. That goal is only sane with a twin. A swarm without a behavioral oracle produces millions of plausible diffs. A swarm with one produces candidates, and the journeys throw most of them away. Evals built backwards is about why the verification side is where effort pays off most.
Evidence-backed PRs
A twin will surface bugs before a swarm finds any optimizations, and that forces a rule for what a PR must show before a person looks at it. From my notes on September 24: "Mark a PR ready for review only after its bug journey passes on the fix, fails on main, and the regression shows no new failures versus main."
Fails on main. A journey that reproduces the bug on current code proves the journey can see the bug. Without this, a pass on the fix means nothing, because a blind test passes everywhere.
Passes on the fix. The obvious half.
No new failures against main. The full journey suite runs on both sides, and the difference must be empty. The comparison is with main, not with an ideal. Known failures on main stay known. Only new ones block.
Before and after videos. These are for people. A reviewer who watches ten seconds of the bug and ten seconds of the fix understands the change faster than one reading the diff.
Production impact. An estimate of how many users hit the journey and what the bug costs in money, with the method written down. This is what decides which fix gets merged first.
The first two rules come from SWE-bench, applied to an app. SWE-bench keeps a task only if it has tests that fail before the reference fix and pass after it, which the paper calls a "fail-to-pass test," and "for each task instance, there is at least one fail-to-pass test which was used to test the reference solution." My regression rule adds the other half: every journey that passed on main must still pass. The difference is that SWE-bench borrows its tests from the repository. Legacy code doesn't have them, so the twin has to write them first.
A verifier earns trust by catching regressions
"Fails on main" is a special case of a broader rule. A verifier's passes mean nothing until you've watched it fail when it should. So I break things on purpose: change a query, remove a guard, invert a condition, and check that some journey goes red. If nothing fails, the suite has a hole there, and every green run over that code was noise. Each journey should carry its own mutation evidence, the list of deliberate breaks it caught.
This is mutation testing, and Google has the best public numbers on running it for real. In State of Mutation Testing at Google (ICSE-SEIP 2018), Goran Petrović and Marko Ivanković describe a system "used by 6,000 engineers in Google on all code changes they author or review, affecting in total more than 13,000 code authors." They evaluated "more than 70'000 diffs, testing 1.1 million mutants and surfacing 150'000 actionable findings during code review." They also say why coverage isn't enough: "coverage alone might be misleading, as in many cases where statements are covered but their consequences not asserted upon." A journey that visits a screen and asserts nothing about it is covered and blind.
The number to watch is how useful developers found the mutants, and it started low. The 2021 follow-up by Petrović, Ivanković, Gordon Fraser and René Just says developers "initially classified 85% of reported mutants as unproductive." The 2018 paper reports usefulness rising "from 20% to 80%" through the feedback loop. The 2021 paper reports productivity going "from 80% to 89%" as feedback was generalized into suppression rules. By then the service was used "by more than 24,000 developers on more than 1,000 projects," across 776,740 changelists and 16,935,148 mutants.
What moved the number was learning what not to mutate. Most early complaints were about mutants in logging, timeouts and configuration flags, which Google calls "arid" lines: places where a surviving mutant is real but a test for it would be pointless. The authors say the effect is hard to measure exactly, but they credit those suppressions with "improvements in productivity from about 15% to 80%." The lesson carries over to a twin. Mutate the code the journeys are supposed to protect, not the log lines, and let reviewer feedback decide which kinds of breaks are worth reporting.
Mutation has a second target that I think is underused: the verifier itself. Delete an assertion from a journey, or make a wait return early, and check that the harness notices. The rules I'd hold any suite to follow from that. Evaluation is deterministic only. On September 7 I wrote (my translation) "I don't want LLM as a judge, the eval should only have deterministic evaluations." Assertions read data (rendered widgets, route traces, backend traffic, database rows) and never prose. Timing tests wait for real load, so a blank page can't win a latency race. Agents may read the app, but the journeys, labels and scorer stay hidden, and the authoritative score comes from outside their sandbox. Show the problem, hide the metric has the argument for hiding the verifier. The point here is narrower: a verifier nobody has tried to fool hasn't earned anything yet.
The same goes for the measurement. When a score comes back at zero, or doubles overnight, the first suspect is the eval, not the code. An implausible result usually means a broken sensor, and a swarm will happily optimize a broken sensor.
Determinism is what makes the oracle cheap
A verifier that gives different answers on the same code isn't an oracle. FoundationDB, TigerBeetle and Antithesis all start by removing nondeterminism.
FoundationDB simulates a whole cluster in one thread. Its testing docs say the simulator can "conduct a deterministic simulation of an entire FoundationDB cluster within a single-threaded process. Determinism is crucial in that it allows perfect repeatability of a simulated run." It "runs tens of thousands of simulations every night," and the team estimates "the equivalent of roughly one trillion CPU-hours of simulation." Their favorite fault, swizzle-clogging, stops the network connections of a random subset of nodes one by one, then restores them in random order. The section ends: "It seems unlikely that we would have been able to build FoundationDB without this technology."
TigerBeetle replays any bug from a seed. Its simulator, the VOPR, stubs out "the clock, network, and disk operations," and because it is "deterministic based on a seed number and the Git commit, we can perfectly reproduce any bugs discovered in testing." It also runs time forward: "One minute of VOPR time is equivalent to days of real-world testing." That's Grieves's front running, applied to a database. TigerBeetle keeps its assertions on in production, on the logic that "it is far better to stop operating than to continue operating in an incorrect state." Back in 2022 one of its developers told HN you could run the simulation yourself by "cloning the repo and running" scripts/vopr.sh.
Antithesis sells determinism as a service. Its docs describe running "multiple copies of your software in a simulation environment that's much more hostile than prod," and add that "because Antithesis' simulation is perfectly deterministic, you get perfect, effortless repro of any problems you find." The pitch is aimed at exactly this problem: "it tests your software without you writing tests — which is vital if you're going to keep up with AI writing code."
A mobile twin can't be as deterministic as FoundationDB, but it can get most of the way. Pin the clock. Seed the database. Fix device models and OS versions. Record and replay the third-party calls you don't control. Every source of randomness you remove turns a flaky red into a real one.
Invariants above the tests
Journeys check behavior from the outside. Some properties are better stated once, as invariants, and checked against every input a generator can find.
AWS writes the design down before the code. Newcombe and coauthors wrote in How Amazon Web Services uses formal methods (2015) that "since 2011, engineers at Amazon Web Services (AWS) have used formal specification and model checking to help solve difficult design problems in critical systems." The full paper explains why: "testing the code is inadequate as a method to find subtle errors in design, as the number of reachable states of the code is astronomical." Its table lists what TLA+ found: two bugs in an S3 fault-tolerant network algorithm, three in DynamoDB's replication and group-membership system, "some requiring traces of 35 steps," and three in EBS volume management. It wasn't a research team's hobby. "Engineers from entry level to Principal have been able to learn TLA+ from scratch and get useful results in 2 to 3 weeks."
Property-based testing finds the cases you didn't write. fast-check defines a property as "an assertion of a relationship between a code's input and output which should hold for all sets of inputs." Its key feature is shrinking: "a failure caused by an array with dozens of elements may result in the developer being shown a failing input with only 2 or 3 items." The model-based mode is the one that matters for agents. You define commands, each with a check that says whether it's allowed in the current state and a run that executes it and asserts. fast-check generates random sequences of commands, shrinks any failing sequence to a minimal one, and gives you a seed and a replay path.
That fits the framework I'm building my personal assistant on. Tardigrade models an agent as "a component tree over an immutable event log, and each component derives a view and enabled transitions as a pure function of the log." If tools are transitions over a log, then a run is a sequence of commands, and fast-check can generate sequences nobody planned. On September 25 I wrote (my translation) that "Tardigrade with fast-check gives a really nice harness abstraction so we can see what we hadn't planned regarding tools."
The invariant I care about most is about money. In my assistant, a tool marked destructive (moving money, writing to an account, creating a public link) refuses to run unless the model sends confirm: true. Tardigrade's host is crash-proof because it re-derives unfinished work from the stored log, which raises a second question: after a crash and a replay, does anything run twice? Here's the shape of the property, written against a plain Harness interface rather than Tardigrade's API:
import fc from "fast-check"
import assert from "node:assert"
type Model = { confirmedTransfers: number }
class Transfer implements fc.Command<Model, Harness> {
constructor(readonly confirm: boolean) {}
check = () => true
run(m: Model, h: Harness) {
h.call("transfer", { amount: 10, confirm: this.confirm })
if (this.confirm) m.confirmedTransfers++
assert.equal(h.executedTransfers(), m.confirmedTransfers)
}
toString = () => `transfer(confirm=${this.confirm})`
}
class Crash implements fc.Command<Model, Harness> {
check = () => true
run(m: Model, h: Harness) {
h.crashAndReplay()
assert.equal(h.executedTransfers(), m.confirmedTransfers)
}
toString = () => "crash"
}
fc.assert(
fc.property(
fc.commands([fc.boolean().map((c) => new Transfer(c)), fc.constant(new Crash())]),
(cmds) => fc.modelRun(() => ({ model: { confirmedTransfers: 0 }, real: newHarness() }), cmds),
),
)
One property holds two invariants: no transfer runs without a confirm, and a crash followed by a replay never pays twice. If either breaks, fast-check returns the shortest sequence that breaks it, something like [transfer(confirm=true), crash], with a seed to reproduce it.
I still have open questions here. Should a violating call be blocked before it executes, or only detected? How do preconditions like "tool B only after tool A" reach the model, instead of living only in the checker? And does writing a TLA+ spec at the start of a project pay for itself at my scale, when AWS reserved it for its hardest systems?
The rest of this layer, types that make wrong states unrepresentable and assertions that run in production, is in Show the problem, hide the metric. One line from that spec belongs here too: "'Crashed scored zero' cannot be constructed."
Structure is not a shortcut
Behavior is one axis. Structure is another, and in September I hoped it would be a shortcut. Sentrux scores a repository with "5 root cause metrics. One continuous score": modularity, acyclicity, depth, equality and redundancy, combined as a geometric mean (Simplicity does not sell has the details). On September 4 my working belief was (my translation) "the higher the sentrux score, the less slop and the easier it is to add things." If that held, structure could be a general reward for any codebase, and a swarm could optimize it.
Any such loop needs two guardrails from the start. A candidate has to pass the behavioral checks before its structure score counts, and nobody gets to reweight the dimensions. The rule I wrote down: "do not allow structural optimization to become a reward-hacking exercise." A note from September 9 was plain about the limit: "sentrux measures shape, not behavior, and its own FAQ says so."
The belief is testable, and the test should be framed so a null would count: "The goal is not to prove that Sentrux works; the goal is to determine, with rigorous statistics, which aspects of repository structure, if any" affect agent success. The measurement definitions get frozen before any statistics run. There are two directions to check. Do stronger models leave higher-scoring code behind? And do higher-scoring repositories make agents succeed more often?
I haven't found a strong public answer. The closest is a May 2026 study from SonarSource, Does Code Cleanliness Affect Coding Agents? The authors built six pairs of repositories that "match on architecture, dependencies, and external behaviour" but differ in static-analysis violations and cognitive complexity, then ran 33 tasks ten times per side, 660 trials, with Claude Code on Sonnet 4.6. Their finding: "code cleanliness does not change the agent's pass rate." At the dataset level the difference was 0.9 percentage points, slightly in favor of the messier side. What changed was cost. On the cleaner side the agent used 7.1% fewer input tokens and 8.5% fewer output tokens, and revisited files it had already edited 33.8% less often. The study is small, it uses a single model, and it comes from a company that sells code-quality tools. It's still the best controlled evidence I know of, and it points one way. Structure seems to change how much an agent spends, not whether it gets the work right.
That changes the order of work. If cleaning a repository doesn't make agents more correct, you can't clean your way to safe agents on legacy code. The twin comes first, because only behavior decides whether a change ships. Structure is a cost optimization you run inside the behavioral gate. A rule I'd written into an agent brief on September 15 applies: "Bias toward rigor over a positive result — a well-supported 'nothing here' is a fine outcome."
Where the twin lies
Three objections are strong enough to change how you build this.
End-to-end tests are flaky and slow. Mike Wacker's Just Say No to More End-to-End Tests (Google Testing Blog, 2015) is the classic case. He follows a composite team that requires 90% of its end-to-end tests to pass before a release and watches the pass rate swing for more than a week, down to 1% when lab hardware failed. The last entry reads: "(Rounds up to 90%, close enough.) No fixes were checked in yesterday, so the tests must have been flaky yesterday." Among what went wrong: "Partner failures and lab failures ruined the test results on multiple days," and "End-to-end tests were flaky at times." A good feedback loop, he argues, is fast, reliable, and isolates failures, and end-to-end tests are weak on all of those. Google's suggestion was "a 70/20/10 split: 70% unit tests, 20% integration tests, and 10% end-to-end tests."
He's right about the pyramid when you have one. Legacy code doesn't. Feathers's essay names the obstacle: "The hardest part is breaking dependencies around a piece of code well enough to be able to exercise it in a test harness." Breaking dependencies is exactly the refactor you're afraid to make without tests. The twin's journeys are the oracle that makes it safe to build the pyramid. Once a module's behavior is pinned at the journey level, an agent can extract it and write the unit tests Wacker wants. Even his own composite story concedes the point: "Despite numerous problems, the tests ultimately did catch real bugs."
Flakiness is the real cost, and the answer is the determinism section above, plus a rule I give every agent: never read an infrastructure failure as evidence about the app. A journey that fails because a simulator host died goes into a separate lane, not into the regression diff.
Twins drift from production. Grieves's model has two links, data flowing from real to virtual and information flowing back. A software twin built once has neither. The backend schema moves, a third-party API changes its error codes, users find a path through the app that no journey takes. Then the twin is green and production is broken. What I'd add is a direct measure of the gap: compare the twin's screen-visit histogram and API traces against sampled production traffic, and treat divergence as a failing check on the twin itself.
A related trap: a twin preserves existing bugs, and that's correct. Feathers tells the story of fixing a bug and then hearing from users who "depended upon the behavior I'd removed." Evidence-backed PRs are how you change behavior on purpose, with the old behavior on record.
Mutation testing is expensive. The 2018 Google paper opens by saying traditional mutation analysis "is computationally prohibitive which hinders its adoption as an industry standard." The 2021 numbers show the scale of the problem and the fix. Even restricted to the lines a changelist touched, traditional mutant generation produced more than 450 mutants for most changelists, with a 25th to 75th percentile range of 460 to 1,734. With arid lines suppressed and one mutant per line, the median fell from 820 to 7. The cost is manageable if you mutate only what changed and only where a test could matter. For a twin, a handful of mutants per PR on the changed code, checked against the journeys that touch it, fits in a budget. Mutating the whole app every night doesn't.
Where this goes
This part is speculation.
If the twin works, it stops being a test environment and becomes a reward function. Every device and journey pair is a deterministic signal, and a swarm searching over versions of the app is optimizing against it without updating any weights. "Millions of versions" is meant literally. Most get thrown out by the regression gate. Some pass it and are faster, and on a mobile app the latency wins multiply across every person who opens it. Find the loops is about choosing which loops are worth closing. A behavior-preserving twin is what makes the app one of them.
The other consequence is about what's valuable. Code generation keeps getting cheaper. A twin of one specific legacy system, with journeys that encode what that system actually does and mutation evidence that they catch breaks, can't be downloaded from a lab. On legacy code, the verifier is the asset and the agents are interchangeable.
Here's where the evidence stops. Each piece has public support: characterization tests from Feathers, mutation testing at Google, deterministic simulation at FoundationDB and TigerBeetle, TLA+ at AWS, and one controlled study on structure and agents. I don't know of anyone who has published a swarm run at millions of versions against a twin of a real app, and I haven't run one myself. Everything past that is a bet.
A PR evidence template
## Change
One sentence: what behavior changes, or "none" for a refactor.
## Journey evidence
- Journey: <id>, path through screens <list>
- On main: FAIL (run <link>, video <link>)
- On this branch: PASS (run <link>, video <link>)
- Refactors: no journey should change; list the journeys that cover the touched code.
## Regression
- Suite: <passed>/<total> journeys, same seed, devices and OS versions as main
- New failures vs main: 0
- Known failures on main, unchanged: <list>
- Infrastructure failures, excluded and rerun: <list>
## Verifier check
- Mutants applied to the changed code: <n>
- Killed by journeys: <n>
- Surviving, with reason: <list>
## Impact
- Users who hit this journey per day (estimate and source): <n>
- Cost of the bug or value of the gain in money (estimate and method): <n>
## Not verified
What this evidence does not cover.
Twin-readiness checklist
Before pointing a swarm at a legacy codebase:
- The whole stack runs locally. Real backend, real app, seeded data, one command to bring it up and one to reset it.
- Journeys are paths on a screen graph. Every screen is a node and every journey is a path, so you can list the transitions nothing covers.
- Coverage follows traffic. The journeys that gate PRs are the most-used ones, weighted by real usage.
- Runs are deterministic. Clock, seeds, devices and third-party responses are pinned. The same commit gives the same result twice in a row.
- Assertions read data, not prose. Rendered elements, routes, backend requests, database rows. No LLM judge.
- Timing waits for real load. A blank screen can't win a latency test.
- Every journey has mutation evidence. At least one deliberate break it catches, recorded next to it.
- The suite itself has been mutated. Remove an assertion or shorten a wait, and confirm the harness notices.
- The verifier is hidden. Agents can read the app, not the journeys, labels or scorer. The authoritative score runs outside their sandbox.
- Infrastructure failures have their own lane. A dead simulator is never evidence about the app.
- Drift is measured. Twin traffic is compared with sampled production traffic on a schedule.
- Impossible scores trigger an investigation. A zero or a sudden doubling is a broken sensor until proven otherwise.
- The PR template is enforced. Nothing is ready for review without fail on main, pass on the fix and an empty regression diff.
Sources
- Michael Feathers, Working Effectively with Legacy Code (2004), quoted in an answer on Software Quality Assurance Stack Exchange, 2017
- Michael Feathers, Characterization Testing, 2016-08-08
- Romano, Vendome, Scanniello, Poshyvanyk, A Multi-Study Investigation Into Dead Code, IEEE TSE, 2018
- Grieves and Vickers, Digital Twin: Mitigating Unpredictable, Undesirable Emergent Behavior in Complex Systems, in Transdisciplinary Perspectives on Complex Systems, Springer, 2017
- Glaessgen and Stargel, The Digital Twin Paradigm for Future NASA and U.S. Air Force Vehicles, NASA, 2012
- Maestro documentation
- Jimenez et al., SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, 2023
- Petrović and Ivanković, State of Mutation Testing at Google, ICSE-SEIP 2018 (DOI)
- Petrović, Ivanković, Fraser, Just, Practical Mutation Testing at Scale: A view from Google, 2021
- FoundationDB, Simulation and Testing
- TigerBeetle, Deterministic Simulation Testing (VOPR); TigerBeetle developer on Hacker News, 2022-03-31
- Antithesis, Welcome to Antithesis
- Newcombe, Rath, Zhang, Munteanu, Brooker, Deardeuff, How Amazon Web Services uses formal methods, 2015 (full text)
- fast-check, What is Property-Based Testing? and Model based testing
- Tardigrade
- sentrux
- Trivedi and Schmitt (SonarSource), Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study, 2026-05-19
- Mike Wacker, Just Say No to More End-to-End Tests, Google Testing Blog, 2015-04-22
- My own notes and agent messages, September 2026 (quotes translated from Portuguese where marked)