Last week I called AI TrustOps the accountability work hiding under a compliance costume. This is the first in that series โ€” nine articles on making AI accountable enough to trust with real decisions. We start where the costume shows up first: the pilot that never becomes production.

Why Most AI Pilots Never Make It to Production โ€” And Why the Model Is Rarely the Reason

Here is a number that should stop every steering committee mid-sentence.

According to RAND Corporation’s research on AI project failure, by some estimates, more than 80 percent of AI projects fail.

For generative AI specifically, MIT’s Project NANDA put an even sharper number on it: 95 percent of organisations are getting zero measurable return on their GenAI investments. It’s worth saying plainly that this finding is still early and has drawn some methodological pushback since it was published โ€” the underlying sample is smaller than the headline suggests, and “zero P&L impact” isn’t the same claim as “the initiative failed.” I’m citing it anyway, because even the more conservative reading of that research points the same direction as RAND’s: the gap between piloted and production-proven AI is large, and it isn’t closing on its own. If anything, that caveat belongs in this article for a reason beyond fairness โ€” an article about the difference between evidence and assumption should not lean on a stat it hasn’t itself interrogated.

The instinctive response to numbers like that is to look at the technology.

Was the model good enough? Was the data clean enough? Was the platform scalable enough? Was the use case too ambitious?

Those questions are not wrong. But they are often asked too late, after the organisation has already fallen in love with the pilot.

RAND’s research points to something more uncomfortable. AI projects do not fail only because the model fails. They fail because the business problem was misunderstood, the wrong metrics were optimised, the solution did not fit the actual workflow, the data foundation was not ready, or the organisation chased the newest technology instead of solving a real institutional problem.

In other words, the model is often where the failure becomes visible. It is rarely where the failure begins.

The Pilot That Works and the Pilot That Scales Are Different Things

Almost every stalled AI initiative has a working pilot somewhere in its history.

The model performed. The demo landed. The steering committee nodded. Someone in the room said, “This is great,” and meant it.

I’ve sat in that meeting. I’ve watched the nodding happen. And I’ve watched, months later, the same initiative quietly still be “in pilot” โ€” not because anything broke, but because nobody in that room was ever going to be the one to ask what happens next.

Because production quietly demands answers that a demo never asks.

Who is accountable if this produces a wrong, costly, or harmful output next week? Not which team built it. Not which vendor supplied it. Not which committee reviewed it. Which named person owns the consequence?

Can someone actually stop it today, without convening a committee? Not in the architecture diagram. Not in the policy document. In practice, the next time it matters.

What is the evidence that this creates value, as opposed to the assumption that it will? A pilot’s success criteria and a production system’s success criteria are rarely the same document. That gap is where many AI initiatives quietly die.

None of these are pure technology questions. They are accountability questions. And pilots are structurally exempt from answering them.

Why “Almost Ready to Scale” Usually Means Something Else

In most organisations, “almost ready to scale” sounds like a technical status. It usually isn’t. And the three reasons it isn’t aren’t peers โ€” one of them causes the other two.

The root cause: nobody owns the kill decision.

Killing a pilot means someone has to stand up and say the evidence isn’t strong enough, or the value isn’t clear enough, or the risk isn’t acceptable yet. That’s an uncomfortable thing to say in a room that applauded the demo. So instead of a decision, the pilot gets silence โ€” and silence defaults to continuation. Everything else downstream follows from this one missing decision.

The first symptom: oversight is real on paper, theoretical in practice.

There is a human in the loop. A review process. A governance committee. A control framework. But when the system produces a bad output, who intervenes, how quickly, with what authority, using which tested mechanism? If nobody had to answer that question in advance โ€” because nobody was forced to decide whether this pilot lives or dies โ€” the answer stays theoretical. Untested oversight isn’t a separate problem from the ownership gap. It’s what an ownership gap looks like from the outside.

The second symptom: success was never defined precisely enough to fail against.

This is the most common failure pattern in enterprise AI, and it too traces back to the same root. A pilot begins with a high-level value hypothesis and enough early promise to keep going. But converting that promise into a hard, measurable, production-grade success standard requires someone to accept the risk that the number comes back negative โ€” which is, again, a kill-decision question wearing a metrics costume. So the standard never gets set. The pilot can’t clearly pass. And because it can’t clearly pass, it also can’t clearly fail. That is how pilots become permanent.

The Real Gap Is Not Between Pilot and Production. The real gap is between excitement and accountability.

A pilot can survive on potential. Production cannot. Production requires named ownership, tested controls, evidence of value, auditability, explainability, operating resilience, escalation paths, data and model clarity, and โ€” underneath all of it โ€” a decision on whether the system should scale, be redesigned, paused, or killed.

That is why many AI pilots do not fail dramatically. They simply lose institutional momentum. They drift from innovation deck to steering committee update to “next quarter” review, technically alive and operationally unresolved. The organisation does not kill them. It just stops believing in them loudly.

What This Means Before Your Next Steering Committee

Before approving the next AI pilot, or letting an existing one sit indefinitely “almost ready,” ask four questions:

Who owns this by name?

Who can stop it in practice?

What evidence proves it works, not merely that it could work?

What would have to be true for us to kill it โ€” and who has the authority to say so?

If those questions feel uncomfortable, that’s the point. They are what separates AI experimentation from AI TrustOps.

Every organisation I’ve seen actually solve this started with that fourth question, not the first three. Naming an owner is easy to write in a RACI chart. Building the actual discipline to answer these four questions before the money is spent โ€” not after โ€” is the harder, rarer capability. It starts one level up from where most of this conversation happens: with the board.

Next in this series: The five questions every board should ask before approving an AI system to scale.


Leave a Reply

Your email address will not be published. Required fields are marked *