In McKinsey’s 2025 survey, 88 percent of respondents said their organisations regularly used AI in at least one business function. Seven percent said AI was fully scaled across the organisation.

That is a scale gap. It does not mean 93 percent of pilots failed.

A pilot can prove that a model performs a task. Production must prove that the organisation can operate the capability reliably. Those are different tests.

The model may summarise a claim, classify a document or recommend an action perfectly inside the demo. The production system still has to find the right data, establish who is asking, respect permissions, fit the workflow, handle timeouts, preserve an audit trail and recover when another system disagrees.

The pilot tested the visible part. The operating environment remained implicit.

What did the pilot actually prove?

A useful pilot answers a narrow question: can this model improve a defined task under controlled conditions?

That evidence matters. It does not establish that the capability can survive contact with production.

Production introduces six boundaries:

  • Data meaning. A field called customer may refer to a policyholder, payer, patient or account owner depending on the system.
  • Identity and access. The model needs a principal, a scope and an expiry. “The application called the API” is not enough.
  • Workflow. A recommendation must reach the right person at the right moment, with a way to disagree.
  • Runtime reliability. External actions can time out after succeeding. A retry can create a duplicate.
  • Audit lineage. Reviewers need to reconstruct which data, transformation, model, policy and human decision produced an outcome.
  • Third parties. Models, data, integration platforms and cloud services introduce separate continuity and transparency risks.

I call the accumulated work across those boundaries the translation tax. It is an analytical model, not an industry statistic. The tax grows when a pilot treats each boundary as somebody else’s later problem.

Legacy systems are doing their job

Legacy infrastructure is often described as technical debt waiting to be removed. That language misses why it remains.

The system closes the books. It pays benefits. It routes care. It controls physical equipment. Organisations trust it because decades of exceptions, rules and recovery procedures have accumulated around it.

The US Internal Revenue Service is an extreme example. The Government Accountability Office has documented core tax systems that rely on assembly language and COBOL, along with rising costs and shortages of relevant skills. A 2025 GAO review reported that replacement of the roughly 60-year-old Individual Master File remained unfinished.

An AI service can sit above that estate. It cannot pretend the estate has disappeared. Every translation must preserve tax-year semantics, taxpayer identity, authority and authoritative-record status.

Healthcare exposes a different problem. Transporting a record through an API does not guarantee that two systems mean the same thing. NHS interoperability guidance separates identifiers, authentication, clinical terminology, medicine codes and exchange formats because each layer can disagree independently. Germany’s DigitalRadar programme reports wide variation in hospital digital maturity and identifies interfaces, compatibility and security as real implementation work.

Operational technology makes the boundary physical. NIST’s guidance for industrial control systems treats safety, reliability and performance as distinct requirements. A cloud-style integration pattern cannot be attached casually to a programmable controller or energy system. A delayed or malformed message can affect equipment and people, not just a dashboard.

None of these examples proves that legacy systems cause most AI failures. They show why integration cannot be reduced to an API call.

Middleware is sometimes the right answer

Replacing the underlying system before using AI sounds cleaner. It can also create a larger failure.

Modernisation can disrupt the service it was meant to improve, so migration, outsourcing and post-launch risk belong inside the operating evidence rather than outside the pilot.

An overlay can therefore be rational transitional architecture. It can expose a limited set of functions, translate old formats and keep authority inside a proven system of record.

The tradeoff needs to be explicit. Each new interface requires authentication, authorisation, encryption, logging, testing, capacity control and lifecycle ownership. The Basel Committee’s digitalisation-of-finance report describes how combining new technology with legacy estates can add complexity, vendor lock-in and cyber exposure.

The mistake is not using middleware. The mistake is budgeting for the model while treating the operating layer as free.

What operating evidence a production pilot reveals

A task test and operating evidence answer different questions.

Decision reconstruction. The evidence follows source data through every transformation, model call, policy check, human intervention and external action. BCBS 239 implementation work continues to identify legacy systems, distributed data and changing lineage as obstacles to end-to-end traceability.

Proposal, permission and outcome. A model recommendation is not authorisation. An attempted action is not confirmed settlement. A production record preserves those states separately, including unknown when the external result cannot be established.

Failure behaviour. Timeouts, stale data, denied permissions, partial dependencies and duplicate requests reveal which operations can retry automatically and which require reconciliation.

Supplier substitution. NIST recommends due diligence, supplier transparency, service-level agreements, monitoring and contingency processes for third-party models, data and systems. Exporting prompts is not the same as recovering evaluations, policies and operating evidence.

Correction after launch. User feedback is a runtime control. Complaints, overrides, appeals, error reports and workflow friction enter a managed process, change priorities and produce measurable follow-up. NIST’s AI Risk Management Framework explicitly calls for feedback, appeal, override, incident response and continual improvement.

A pilot that cannot answer these questions may still be a valuable experiment. It has not yet proved production readiness.

The model was only one dependency

The model demo receives attention because it gives visible outcomes quickly. The translation tax appears later in data mapping, access reviews, incident procedures, service contracts and arguments about which record is authoritative.

That work is slower. It is also where the organisation decides whether AI remains a feature or becomes an operating capability.

Your pilot may have shown that the model works. The next test is whether the surrounding system can carry the consequence.