A Rewrite Passes Every Test and Still Changes the Business
Picture a claims system rebuilt from a thirty-year-old mainframe application. The new service passes its acceptance tests, the cutover goes smoothly, and three weeks later finance notices that a class of policies is no longer being rounded the way regulators expect. Nobody removed that rule on purpose. It was a four-line conditional added during a 1998 audit, never written into any requirements document, and the team writing the new service never saw it because no requirement pointed at it.
This is the part of modernization that estimates tend to miss. The code is the only complete record of what the business actually does, including rules that were patched in after incidents, regulatory changes, and one-off customer deals. The people who remember why those rules exist retire or move on. Secondary estimates put the average COBOL programmer at about 55 years old with roughly 10 percent of that workforce retiring each year. Those figures are repeated more than they are measured, but the direction is not in dispute, and the GAO report points the same way on skills.
Two 2026 research papers show what is newly possible. One describes a seven-agent pipeline that discovered 280 business journeys and extracted 1,612 business rules across 39 repositories. Another, called Reversa, tags each recovered claim as confirmed, inferred, or gap, and links it back to the code that supports it. Both are early work with real limits, and we cover them below. The practical lesson is the same in each: extraction is now fast, so the discipline has to move to verification and ownership.
Reading legacy code is no longer the slow part. Deciding which of its rules still belong to the business is.
Six Kinds of Hidden Rules, Ranked by How Badly They Hurt When Lost
| Hidden rule type | Where it lives and how it gets lost | Severity |
|---|---|---|
| Regulatory and audit patches | Conditionals added after an audit finding or rule change, with no ticket history, so a rewrite drops them silently | Critical |
| Rounding, ordering, and boundary behavior | Implicit choices in arithmetic and sort order that downstream reports and reconciliations depend on | Critical |
| Rules in stored procedures and job schedulers | Logic outside the application repository, which static code analysis does not see | High |
| Customer and contract exceptions | Hard-coded account IDs or product codes that encode a deal struck years ago | High |
| Error handling and retry behavior | Fallbacks that keep batch runs alive, which the new design replaces with different failure modes | Moderate |
| Dead branches that look live | Code for products or regions that no longer exist, which gets migrated and tested at real cost | Lower |
Not sure where your legacy business rule gaps are?
10decoders reviews how one legacy system is documented, tested, and owned, from stored procedures to the people who still know why a rule exists. You get a ranked list of what is most likely to be lost in a rewrite and a plan to recover it first.
Book a Free AI Assessment →What the Research Supports and Where It Stops
The seven-agent pipeline is the more ambitious of the two. It reports generating artifacts for 18 services in an average of 8.78 minutes each, and it found 115 extracted rules that were independently corroborated against production defect triage records. That corroboration step is the most useful idea in the paper, because it ties a rule to evidence of real behavior. The authors are also direct about limits. The manual baseline they compare against, 117 person-hours per service, is modeled and not measured. Only one Java codebase was evaluated. The static-first method misses behavior that is resolved at runtime, such as reflection, dependency injection, and stored procedures. High-confidence extraction also needs usable test suites, which many legacy systems do not have.
Reversa takes a narrower path. In an exploratory COBOL-to-Go study of an educational ATM system, it classified 517 claims into 490 confirmed, 24 inferred, and 3 gaps, and it kept the gaps visible for a human to resolve instead of filling them with a guess. The authors describe the result as exploratory evidence that the approach is plausible. There was no independent expert audit and no controlled baseline, and the final parity validation was incomplete. Treat both papers as design patterns, not as benchmarks you can quote in a business case.
For an enterprise team, the transferable pattern is three habits. Link every recovered rule to the code or test that supports it. Label each one as confirmed, inferred, or unknown. Send the inferred and unknown ones to a named business owner. An agent that can say “I could not determine this” is far more useful than one that always produces a clean-looking specification.
Rules Live in Heads and Source Files
The rewrite team works from screens, a few documents, and interviews with whoever is left. Missing rules surface as production incidents after cutover.
Agents Draft Rules, Humans Check Them
Static analysis and agents produce rule candidates tagged confirmed, inferred, or gap. Reviewers work the uncertain ones, but sign-off is informal and the rules sit in a document.
Rules Become Tests With Named Owners
Each rule has an owner, evidence, and a parity test that runs against both old and new systems. Retiring a module requires its rules to pass.
Checklist: Is Your Rule Recovery Ready to Support a Cutover?
Choose one legacy system you plan to modernize and check each item. Anything you cannot show as a register entry, a test, or a named person is a gap that will reach production.
The goal of an extraction run is a shorter list of questions for the business, not a longer document.
What to Do This Week
1. Pick one module and list where its logic actually lives
Choose the legacy module you plan to touch first. Write down every place its behavior is defined: source files, stored procedures, scheduler definitions, copybooks, config tables, and spreadsheets that operations teams keep on the side. Ask two long-tenured people to read your list and add what is missing. The omissions they name are your first estimate of what an automated reading would not see.
2. Run a bounded extraction and keep the labels
Point an extraction tool or agent at that one module, not the whole estate. Require that every output rule cites its source and carries a label of confirmed, inferred, or gap. Count how many land in each bucket. A high share of gaps is useful information about where documentation and tests are thinnest, and it tells you how much interview time the rewrite really needs.
3. Match the rules against your incident history
Take the last two years of production defects and incidents for the module and check which extracted rules explain them. Rules that map to real failures go to the top of the review queue. Rules that match nothing go to a business owner with one question: is this still needed? Expect some of those to turn out to be dead branches you can skip migrating.
4. Turn five critical rules into parity tests
Select the five rules with the highest financial or regulatory weight and write each as a test with boundary inputs, such as rounding on the last cent, date edges, and empty fields. Run them against the legacy system today to capture expected outputs, then keep them as the acceptance gate for the replacement. If you cannot produce a test for a rule, that is a sign the rule is not understood yet.
Let 10decoders Recover the Business Rules in Your Legacy System Before You Rewrite It
We take one system and map where its logic lives, run a labeled rule extraction, check the results against your incident history, and draft parity tests for the highest-risk rules. You leave with a ranked rule register and a retirement gate your business owners can sign.
