ONGOING INDEPENDENT COVERAGE

Questions Buyers Are Asking# Can Agentic AI Actually Work on Top of a System Nobody Fully Understands Anymore?

The legacy modernization market itself is growing fast on the back of this pressure — from roughly $29 billion in 2026 to over $66 billion by 2031.

Last Update: 6 September 2026

Most large companies want AI. Most large companies also run old, tangled systems that nobody currently on staff fully understands — the person who wrote the logic left years ago, and the documentation, if it ever existed, didn't keep up. Recent research puts real numbers on this problem: developers spend roughly 58% of their time just reading old code, versus 5% writing new code, and one widely cited study found 60% of AI leaders name legacy systems as the primary barrier to deploying agentic AI at all. The legacy modernization market itself is growing fast on the back of this pressure — from roughly $29 billion in 2026 to over $66 billion by 2031. Thoughtworks has built a specific, named product for exactly this intersection. The fair question is how far "understanding the old system automatically" actually goes before a human still has to step in.


How big this pressure is getting, in plain numbers


Figure

What it means

Legacy modernization market, 2026

~$29 billion

Already a large, active spending category

Legacy modernization market, 2031 (forecast)

~$66 billion

More than doubling in five years

Time developers spend reading old code vs. writing new code

~58% vs. ~5%

Understanding the old system is the actual bottleneck, not writing new code

AI leaders naming legacy systems as the top barrier to agentic AI

~60%

This is a widely shared, not rare, problem


The three shapes of the problem, side by side


Understanding the old system

Turning that understanding into a spec

Building and maintaining the new thing

What it answers

What does this old code actually do, and why

What should the replacement or extension do, precisely

Who builds it, tests it, and keeps it working as things change

Historically done by

A person reading code for weeks or months

A business analyst or architect, by hand

Developers, by hand, over months or years

Biggest risk if skipped

Nobody knows what they're actually replacing

The new system misses rules nobody remembered to write down

The new system drifts out of date again, the same way the old one did

Sector specialty worth naming: teams modernizing a finance or operations backbone

This problem sits inside almost any large company running an old core system underneath a modern front end — banking, insurance, manufacturing, retail, it doesn't matter which. The company worth naming here is Thoughtworks. In January 2026 it launched a named, specific product for this exact problem: AI/works, an agentic development platform built to read and interpret old systems automatically, turn that understanding into structured specifications enriched with regulatory and security context, and then use agentic workflows to generate the code, tests, and deployment pipelines that follow.

1. Does it actually understand the old system, or summarize it plausibly?

There's a real difference between an AI system that correctly reconstructs what old code does and one that produces a confident-sounding but subtly wrong summary — and the second kind is far more dangerous in a regulated, high-stakes system.

My reading: Thoughtworks describes AI/works as using AI-enabled reverse engineering to interpret legacy applications and convert that into structured specifications. That's a real, specific claim — reverse engineering implies working backward from actual code behavior, not just describing it in general terms. What's not public is how that output gets checked for accuracy before anyone builds on top of it.

Ask what the verification step looks like: who checks the AI's understanding of the old system against reality before it becomes the basis for new code?

2. Are the "months instead of years" numbers real, or a general claim?

This is the single most quotable claim in Thoughtworks's own announcement, and also the one with the least specific evidence behind it.

My reading: as of 6 September 2026, Thoughtworks's own public statements describe modernization work moving from years to months, with cost reductions and higher code quality — but no named clients and no specific figures were given beyond that general comparison. That's a real claim worth taking seriously, and also one that's fair to ask to see backed with a number.

Ask for one named or anonymized client example with an actual before-and-after timeline in months, not the general "years to months" comparison.

3. What happens when the business rules change after the new system is live?

Old systems accumulate small, undocumented business rules over years. A modernization approach that misses this ongoing part just recreates the same problem on a shorter timescale.

My reading: Thoughtworks describes an ongoing approach after deployment — the platform is said to regenerate affected components automatically when requirements evolve, specifically to avoid falling back into manual patching. This is a genuinely differentiated, specific claim worth crediting.

Ask for a real example of a requirement changing after go-live, and exactly what regenerated automatically versus what still needed a person.

4. Does it work the same way on a regulated system as an ordinary one?

Not every legacy system carries the same weight. A retail inventory system and a core banking ledger are not equally forgiving of an AI system getting something subtly wrong.

My reading: Thoughtworks states the specifications are enriched with regulatory, security, and industry context — a real and relevant claim for regulated industries specifically. What's less visible is whether that context-enrichment has been proven out in a genuinely regulated environment, or whether it's a stated design intention still being tested at scale.

Ask for a specific example in a regulated industry — banking, insurance, healthcare — not just a general claim that regulatory context is included.

5. How does the 90-day delivery promise actually hold up on a large, tangled system?

A "3-3-3" delivery model — idea to production in 90 days — is an aggressive, specific promise.

Our take: a bold, specific number is more useful to a buyer than a vague promise, because it can actually be checked against reality. The fair test isn't whether it's achievable on a clean, well-scoped pilot — it's whether it holds on the messiest, most tangled system in your own estate, not the easiest one chosen to showcase the platform.

Ask which of their own past projects used the full 90-day model end to end, and how large or tangled that system actually was.

Where it fits

A large organization with real, painful legacy systems, wanting to move faster than a traditional multi-year rewrite would allow, with a genuine hybrid estate of old and new technology.

Where it does not fit

A company expecting to skip human verification of the AI's understanding of a regulated, high-stakes system, or one hoping for the 90-day timeline on the single most tangled system in their entire estate without first testing on something smaller.

FAQs

Is this just rebranded outsourced development with AI added on? Based on public material, no — the specific mechanism (reverse engineering old code into structured specs, then agentic workflows building from those specs) is a genuinely different approach from traditional staff-augmented rewrites, though it's still new enough that independent, large-scale proof is limited.

Is Thoughtworks the only company doing this? No — several major firms have entered agentic development generally, by Thoughtworks's own account. What's more specific to this piece is the legacy-first framing, rather than a tool built mainly for writing new code.

What's the one thing most buyers forget to check? Whether "understood by AI" has been verified by a human against the actual old system before anything gets built on top of that understanding — skipping this step is where the real risk sits.



This is a piece of opinion — our reading of what buyers should ask, based on public material available as of the date noted above. It is not a statement of fact about any company. No company mentioned pays for the mention. Any company named here can write to hello@analystlayer.com; we respond within three working days and update the piece where the input is factual, with the update dated on this page. Another version of the analysis can be found at here.