The Demo Is a Question, Not an Answer
Ninety-five percent of AI pilots are said to fail. They fail because companies invest before they prove adoption. Reverse the order, and let the demo find the truth instead of selling a future.
Everyone is quoting the same number. Ninety-five percent of corporate AI pilots fail. It comes from an MIT study that went off like a firework in the late summer of 2025, and it has been borrowed to argue for an AI bubble, an AI winter, and every mood in between. Before you use it, read what it actually measured.
MIT’s Project NANDA found that only about five percent of enterprise AI pilots produced rapid revenue acceleration, and the rest delivered little to no measurable impact on the bottom line (as Fortune reported it). That is a finding about return, not about whether the technology worked. And MIT’s own explanation for why is the part almost nobody quotes: not the quality of the models, but a learning gap, the failure of tools and organizations to adapt to one another. The models were fine. The institutions never changed around them.
You do not need a study to recognize the shape of it. Julie Averill, eight years the chief information officer of Lululemon, described it firsthand in the New York Times: tools that dazzled in a demo and stalled the moment they met a real company’s mismatched data, tangled systems, and decades of human exceptions. She names the belief that a powerful model lets you skip that work AI wishing, and its cousin, describing the outcome you hope for as if it had already arrived, AI washing.
Put those together and the ninety-five percent stops looking like a technology problem. It is a sequencing problem. The pilots that fail are the ones a company invests in before it has proven anyone will adopt them. And the demo is what does the damage, because a good demo is a comfortable way to lie to yourself. It proves something is possible, and possibility gets mistaken for a reason to build.
So invert it. The demo is a question, not an answer. Its job is not to sell a future. Its job is to find out, fast and cheap, what the real problem is and whether anyone actually wants it solved. Prove the pull first. Only then do you spend.
The receipts
This is not a thought experiment. At Avidity I ran pilots this way on purpose, and the stalled kind became rare. The full engagement is its own case study. Here is the transferable shape of it, in four moves.
Buy what you can. The same MIT data found that buying tools from specialized vendors succeeded roughly three times as often as building them in-house. So the first gate is buy versus build, and only the work you cannot or will not buy should ever reach a build conversation at all. Most asks end here, cheaply, with something you did not have to make.
Put a person at intake, not a queue. An AI Product Partner takes the request and first asks whether a no-code tool already solves it. This is the role doing exactly the unglamorous work the ninety-five percent skipped: reframing the ask before anyone writes a requirement, because the person asking usually cannot yet name what they need and tries to solution instead of describing the pain.
Climb, do not leap. When something does have to be built, it earns each rung by proving traction on the one below, and each rung is a real, named thing. A Sub24 is a clickable prototype an engineer builds in under a day, whose only job is to show the person what they actually asked for. An App is that single workflow built for real and put in front of them to use. A Program is what an App becomes once demand is proven and it earns executive sponsorship and a budget. A Platform is the rare survivor that graduates into durable, shared infrastructure the other rungs run on. This is the Scope Ladder applied to a build, one order of magnitude at a time, and it is what keeps the demo honest: a Sub24 surfaces the real requirement in a day instead of winning funding for a roadmap.
Keep the right to stop. Because investment tracks traction, an ask that solves its problem early is a success that ends, not a plan that has to justify itself. A request from Legal to compare contract clauses across many contracts had a full platform plan on paper. We stopped it at the program layer, because the pain was gone sooner than we expected, and the platform we never built is budget that went to the next real problem.
The learning gap is closed by people
None of this is only a technology discipline. Every rung was paired with enabling and training the people who would use it, so traction meant genuine adoption and not a login count. That is the whole point of the MIT finding, read honestly. The gap that sinks the ninety-five percent is a learning gap, between a tool and an organization that never taught itself to use it. You close that gap deliberately, with people, or you do not close it at all. A pilot is not the software arriving. It is the institution changing to meet it.
Monday morning
Take your loudest AI pilot, the one with a champion and a roadmap. Before you approve its next dollar, answer the question a Sub24 would force on you: what have you actually proven that someone will use, as opposed to what you have proven is possible? If the honest answer is that you have a great demo, you do not have a pilot yet. You have a question you have not asked cheaply enough. Ask it first. Let the answer decide what you build.
Cheers,
-Titus
Get it in your inbox.
New Issues, FAQs, and Case Studies as they go out. Each one names something, explains something, or hands you something you can use on Monday. Subscribe, and I will send each as it goes out.
Prefer the tool you already think in? Here is how to read it in your chatbot.