Why this matters now: In June 2026, Gartner predicted that more than 70 percent of mainframe exit projects started this year will fail to deliver their intended benefits, mainly because buyers overestimate what generative AI can do to legacy code. Gartner also expects three quarters of the vendors in the mainframe exit market to change their business model or close by 2030. Budgets are being approved on the strength of demos, and the demos rarely show the testing.

Can the Tool That Explains the Code Also Prove the Migration?

A team feeds a forty-year-old COBOL program into an AI assistant and gets back a clear summary within minutes. The summary names the business rules, flags dead code, and even proposes a Java version. The demo is convincing, and the steering committee reads the result as a schedule cut from years to months.

The trouble is that two different claims are being treated as one. The first claim is that AI can understand legacy code. The evidence for that is good. IBM reports that a national social insurance organization cut the time needed to analyze and locate superfluous COBOL by up to 94 percent using watsonx Code Assistant for Z, and the Open Mainframe Project describes discovery, documentation, and business rule extraction as the areas where AI tooling is strongest today. The second claim is that AI can safely convert that code to another language and platform while preserving every behavior. That claim is much weaker, and it is the one the business case usually rests on.

Behavior in an old system is broader than the logic you can read. It includes how packed decimals round, how a date field behaves at a month boundary, what happens when a batch job is restarted halfway, and which workaround an operator has run every Friday since 2003. Translated code can look right, pass a compile, and still return a different number on one input in ten thousand. In a payments, claims, or policy system, that one input is the whole risk. Modernization work fails at the point where nobody can prove equivalence, not at the point where nobody can read the code.

AI made reading legacy code cheap. It did not make proving the new code identical any cheaper.
70%+
Of mainframe exit projects initiated in 2026 will fail to deliver intended benefits, tied to overestimating generative AI. Source: Gartner press release, June 18, 2026.
75%
Of mainframe exit vendors will pivot their business model or close by 2030. Source: Gartner press release, June 18, 2026.
3 proofs
Discovery, equivalence, and reversibility. The three things 10decoders asks a team to show before approving any legacy code cutover.

Where AI-Assisted Mainframe Migrations Lose Behavior

Failure pointWhat closes itSeverity
Translated code is accepted because it compiles and reads wellRun captured production inputs through old and new code and compare outputs field by fieldCritical
Numeric and date semantics differ between COBOL and the target languageExplicit type mapping with tests for rounding, overflow, and boundary datesCritical
Business rules live only in the heads of retiring engineersAI-assisted rule extraction reviewed and signed off by business ownersHigh
Batch and throughput behavior is never baselinedMeasured performance targets for peak load and batch windows before designHigh
Data migration is planned as an afterthought to code conversionSeparate workstream and acceptance gate for data, with reconciliation reportsModerate
No tested path back to the legacy systemStaged cutover by business function with a rehearsed rollbackLower

Not sure where your mainframe migration plan would break?

10decoders reviews your migration approach against the behavior your legacy systems actually have today. You get a ranked list of equivalence, data, and rollback gaps and a scoped plan to close them.

Book a Free AI Assessment →

Testing Is the Bulk of the Work, Not the Final Step

One quality engineering practice describes a supermarket chain migrating an RPG system where more than 90 percent of the project was testing. Validation took around 18 months, while the code changes themselves took about a month and a half. It is a single case, but it matches what most engineers who have done a real migration will tell you. The conversion is the small part, and the evidence that the conversion is correct is the large part.

The technique that makes this manageable is characterization testing. You do not start from a specification, because the specification is usually out of date. You start from the running system. Record real inputs and the exact outputs the legacy system produces, then treat those pairs as the definition of correct. When the migrated code disagrees with a recorded output, someone decides whether that is a defect or an old bug that the business is happy to retire. Either way, the decision is made with evidence, and it is written down.

This is also why Gartner points buyers toward a platform-smart strategy rather than a single exit path. For large mainframe estates, it suggests modernizing within the existing environment and using generative AI for understanding and refactoring. For the largest group of medium environments, it suggests optimizing what stays and exiting selectively, case by case. IBM adds a useful point about scope: roughly 40 percent of COBOL runs on distributed platforms rather than on mainframes, so the language is not the platform, and translating syntax is not the same as replacing a runtime, a data architecture, and transaction integrity.

Stage 1
Where most teams start

Translation by Demo

An AI tool converts sample programs and the results look right. There is no recorded baseline, so correctness is judged by reading the output.

Stage 2
Where most teams stall

Tested Against Sample Data

Unit tests exist for the converted modules, mostly written from the new code. Edge cases from production and batch behavior are still untested.

Stage 3
Where mature teams operate

Proven Against the Legacy System

Production inputs and outputs are the oracle, a parallel run has numeric exit criteria, and every cutover step has a rehearsed way back.

Behavioral Equivalence Checklist for Legacy Code Cutover

Run this against every program or service group before it moves off the legacy platform.

Checks to pass before cutover

Every business rule has a written sourceEach rule in the legacy code is extracted, described in plain language, and signed off by someone in the business, not just by the engineer who read the code.
Outputs are captured before any code changesReal production inputs and their exact outputs are recorded, including rounding, date handling, and error paths, so the old system becomes the test oracle.
Data types are mapped explicitlyPacked decimals, fixed-width records, and character encodings each have a documented target type and a test that proves no precision is lost.
A parallel run has defined exit criteriaOld and new systems process the same live volume for a set period, with a numeric mismatch threshold agreed before the run starts.
Batch windows and throughput are measuredPeak load, batch duration, and transaction latency are baselined on the current platform and compared with the target, not assumed to match.
Every cutover step can be reversedEach release has a tested way back to the legacy path, including data written by the new system during the trial.
Tacit knowledge is captured while people are still hereRetiring engineers record the odd behaviors, manual workarounds, and reasons behind them, and the notes are linked to the code.
Language conversion is separated from platform changeCode translation, data migration, and runtime replacement are planned and tested as three workstreams with separate acceptance gates.
The migration is finished when you can show the old and new systems agree, not when the new code runs.

What to Do This Week

01 Pick one program and record its real inputs and outputs

Choose a batch job or transaction that moves money or changes a customer record. Capture a week of production inputs with the outputs the current system produced, including rejected records and error messages. Store them somewhere versioned. This set of pairs is your first characterization suite, and it costs far less than the first defect it will catch.

02 Use AI for discovery, and have a business owner check every rule

Run your assistant against that one program and ask for the business rules in plain language. Then put the list in front of the person who owns the process and ask them to mark each rule correct, wrong, or unknown. The unknown column is your risk register, and it is usually longer than anyone expected.

03 Write down the numeric mapping before any code is converted

List every packed decimal, signed field, and date format in the program, then note the target type and rounding rule for each. Add a test for the largest value, the smallest value, a zero, and a month-end date. Most silent conversion defects sit in this small table.

04 Agree the parallel run exit criteria before you run anything

Decide how long old and new systems will process the same volume, what mismatch rate is acceptable, and who can approve the cutover or stop it. Put the numbers in writing and add the rollback step. If those criteria cannot be agreed this week, the program is not ready to migrate, and that is worth knowing now.

Let 10decoders Prove Your Mainframe Migration Before You Cut Over

We review your legacy estate, extract the business rules that matter, build the characterization tests, and set parallel run and rollback criteria, so your modernization plan rests on evidence rather than a demo.