Why this matters now: The US Government Accountability Office found that ten critical federal legacy systems, aged from about 8 to 51 years, cost roughly $337 million a year to operate and maintain, and it links the use of languages like COBOL to a shrinking pool of people with the right skills. McKinsey’s estimate, widely repeated, is that around 70 percent of software at Fortune 500 companies is more than two decades old. Agentic modernization tools have made reading that code cheap for the first time, which makes the real bottleneck visible: deciding which of the rules they find are still the business.

A Rewrite Passes Every Test and Still Changes the Business

Picture a claims system rebuilt from a thirty-year-old mainframe application. The new service passes its acceptance tests, the cutover goes smoothly, and three weeks later finance notices that a class of policies is no longer being rounded the way regulators expect. Nobody removed that rule on purpose. It was a four-line conditional added during a 1998 audit, never written into any requirements document, and the team writing the new service never saw it because no requirement pointed at it.

This is the part of modernization that estimates tend to miss. The code is the only complete record of what the business actually does, including rules that were patched in after incidents, regulatory changes, and one-off customer deals. The people who remember why those rules exist retire or move on. Secondary estimates put the average COBOL programmer at about 55 years old with roughly 10 percent of that workforce retiring each year. Those figures are repeated more than they are measured, but the direction is not in dispute, and the GAO report points the same way on skills.

Two 2026 research papers show what is newly possible. One describes a seven-agent pipeline that discovered 280 business journeys and extracted 1,612 business rules across 39 repositories. Another, called Reversa, tags each recovered claim as confirmed, inferred, or gap, and links it back to the code that supports it. Both are early work with real limits, and we cover them below. The practical lesson is the same in each: extraction is now fast, so the discipline has to move to verification and ownership.

Reading legacy code is no longer the slow part. Deciding which of its rules still belong to the business is.
$337M
Annual cost to operate and maintain ten critical federal legacy systems, aged about 8 to 51 years. Source: US GAO, report GAO-23-106821.
~70%
Share of Fortune 500 software estimated to be over two decades old. Source: McKinsey, as cited in Pragmatic Coders’ legacy code statistics.
1 register
One business-owner-signed rule register per system before any module is retired. This is working guidance from 10decoders delivery practice, not a survey figure.

Six Kinds of Hidden Rules, Ranked by How Badly They Hurt When Lost

Hidden rule typeWhere it lives and how it gets lostSeverity
Regulatory and audit patchesConditionals added after an audit finding or rule change, with no ticket history, so a rewrite drops them silentlyCritical
Rounding, ordering, and boundary behaviorImplicit choices in arithmetic and sort order that downstream reports and reconciliations depend onCritical
Rules in stored procedures and job schedulersLogic outside the application repository, which static code analysis does not seeHigh
Customer and contract exceptionsHard-coded account IDs or product codes that encode a deal struck years agoHigh
Error handling and retry behaviorFallbacks that keep batch runs alive, which the new design replaces with different failure modesModerate
Dead branches that look liveCode for products or regions that no longer exist, which gets migrated and tested at real costLower

Not sure where your legacy business rule gaps are?

10decoders reviews how one legacy system is documented, tested, and owned, from stored procedures to the people who still know why a rule exists. You get a ranked list of what is most likely to be lost in a rewrite and a plan to recover it first.

Book a Free AI Assessment →

What the Research Supports and Where It Stops

The seven-agent pipeline is the more ambitious of the two. It reports generating artifacts for 18 services in an average of 8.78 minutes each, and it found 115 extracted rules that were independently corroborated against production defect triage records. That corroboration step is the most useful idea in the paper, because it ties a rule to evidence of real behavior. The authors are also direct about limits. The manual baseline they compare against, 117 person-hours per service, is modeled and not measured. Only one Java codebase was evaluated. The static-first method misses behavior that is resolved at runtime, such as reflection, dependency injection, and stored procedures. High-confidence extraction also needs usable test suites, which many legacy systems do not have.

Reversa takes a narrower path. In an exploratory COBOL-to-Go study of an educational ATM system, it classified 517 claims into 490 confirmed, 24 inferred, and 3 gaps, and it kept the gaps visible for a human to resolve instead of filling them with a guess. The authors describe the result as exploratory evidence that the approach is plausible. There was no independent expert audit and no controlled baseline, and the final parity validation was incomplete. Treat both papers as design patterns, not as benchmarks you can quote in a business case.

For an enterprise team, the transferable pattern is three habits. Link every recovered rule to the code or test that supports it. Label each one as confirmed, inferred, or unknown. Send the inferred and unknown ones to a named business owner. An agent that can say “I could not determine this” is far more useful than one that always produces a clean-looking specification.

Stage 1
Code is the spec

Rules Live in Heads and Source Files

The rewrite team works from screens, a few documents, and interviews with whoever is left. Missing rules surface as production incidents after cutover.

Stage 2
Extracted and labeled

Agents Draft Rules, Humans Check Them

Static analysis and agents produce rule candidates tagged confirmed, inferred, or gap. Reviewers work the uncertain ones, but sign-off is informal and the rules sit in a document.

Stage 3
Owned and executable

Rules Become Tests With Named Owners

Each rule has an owner, evidence, and a parity test that runs against both old and new systems. Retiring a module requires its rules to pass.

Checklist: Is Your Rule Recovery Ready to Support a Cutover?

Choose one legacy system you plan to modernize and check each item. Anything you cannot show as a register entry, a test, or a named person is a gap that will reach production.

What a defensible legacy rule recovery includes
Scope includes everything that runs, not only the main repositoryStored procedures, job schedulers, batch control files, and configuration are inventoried alongside application code, because that is where static analysis goes blind.
Every rule points to evidenceEach register entry cites the file and line, the test, or the production record that supports it. A rule with no evidence is a guess and is labeled as one.
Confidence is explicitRules are tagged confirmed, inferred, or gap. Gaps stay visible and are never merged into confident prose by the tool or the writer.
A business owner signs each critical ruleRegulatory, financial, and customer-facing rules have a named person who confirms whether the rule is still wanted, changed, or retired. Engineers do not decide that alone.
Rules are checked against real defect historyExtracted rules are compared with past incidents and defect tickets. A rule that explains a real defect gets higher priority than one that appears nowhere.
Each critical rule becomes a parity testThe same inputs run through the old and new system, including boundary values and rounding cases, and the outputs are compared automatically on every build.
Runtime behavior is observed, not only readProduction traces or shadow runs confirm what the code actually does, since reflection, dependency injection, and database logic hide behavior from static reading.
Retirement is gated by rule coverageNo legacy module is switched off until its register entries are owned, tested, and passing. Dead branches are retired deliberately, with the owner’s agreement.
The goal of an extraction run is a shorter list of questions for the business, not a longer document.

What to Do This Week

1. Pick one module and list where its logic actually lives

Choose the legacy module you plan to touch first. Write down every place its behavior is defined: source files, stored procedures, scheduler definitions, copybooks, config tables, and spreadsheets that operations teams keep on the side. Ask two long-tenured people to read your list and add what is missing. The omissions they name are your first estimate of what an automated reading would not see.

2. Run a bounded extraction and keep the labels

Point an extraction tool or agent at that one module, not the whole estate. Require that every output rule cites its source and carries a label of confirmed, inferred, or gap. Count how many land in each bucket. A high share of gaps is useful information about where documentation and tests are thinnest, and it tells you how much interview time the rewrite really needs.

3. Match the rules against your incident history

Take the last two years of production defects and incidents for the module and check which extracted rules explain them. Rules that map to real failures go to the top of the review queue. Rules that match nothing go to a business owner with one question: is this still needed? Expect some of those to turn out to be dead branches you can skip migrating.

4. Turn five critical rules into parity tests

Select the five rules with the highest financial or regulatory weight and write each as a test with boundary inputs, such as rounding on the last cent, date edges, and empty fields. Run them against the legacy system today to capture expected outputs, then keep them as the acceptance gate for the replacement. If you cannot produce a test for a rule, that is a sign the rule is not understood yet.

Let 10decoders Recover the Business Rules in Your Legacy System Before You Rewrite It

We take one system and map where its logic lives, run a labeled rule extraction, check the results against your incident history, and draft parity tests for the highest-risk rules. You leave with a ranked rule register and a retirement gate your business owners can sign.