Why Most AI Pilots Never Pay Off
Most enterprise AI pilots demo well and deliver nothing. Here is what the failure data shows, why budgets overrun, and what the small minority do differently.

Table of contents
Someone in your company ran an AI pilot this year. It demoed well. There were slides, a promising accuracy number, and a room of people nodding.
Then it stopped. Not cancelled exactly, just never quite rolled out, still sitting in a staging environment while the budget quietly moves elsewhere.
That outcome is not unusual. It is the norm. Research on enterprise generative AI found that roughly 95% of pilots never produce measurable impact on profit and loss. Adoption is close to universal and returns are not. Understanding the gap between those two facts is now one of the more valuable things a manager can do.
Key Takeaways
- About 95% of enterprise generative AI pilots deliver no measurable P&L impact
- 79% of enterprises reported AI cost overruns in the past 12 months
- Adoption sits near 88%, but only 39% report measurable EBIT impact and roughly 6% attribute more than 5% of EBIT to AI
- Pilots rarely fail on model quality. They fail on workflow, data access and unclear ownership
- Spending is forecast at $2.59 trillion for the year, up 47%, so the waste is being funded regardless
The number every budget holder should know
Start with the spread between adoption and return, because the two get confused constantly.
Stanford research puts organisational adoption around 88%, and Deloitte's State of AI in the Enterprise tracks a similar picture. Nearly every company of any size is doing something. But only 39% report measurable impact on earnings before interest and tax, and roughly 6% qualify as genuine high performers, meaning they attribute more than 5% of EBIT to AI.
So the picture is not that AI does not work. Around one in sixteen companies is getting serious money out of it. The picture is that most of the spending is not reaching that outcome, and the gap between the 6% and everyone else is where the useful lessons sit.
Worth noting who is spending hardest. Financial services firms average about $3,200 per employee on AI, roughly 2.6 times the cross-industry average. Concentration of spend has not produced a matching concentration of returns.
Why an AI pilot fails, and it is rarely the model
Post-mortems tend to blame the technology. The pattern in the data points somewhere else.
The model is usually the part that worked. A pilot reaches a demo precisely because the model performed well enough on a clean, narrow slice of the problem. What breaks is everything around it.
Three failure modes recur.
- No workflow home. The tool exists beside the job instead of inside it. If a support agent has to leave their ticketing system, paste text into another window and copy an answer back, the tool costs time on a busy day and gets abandoned
- Data that was fine for a demo and not for production. Pilots run on an extracted, cleaned sample. Production runs on live systems with permissions, gaps and records nobody has maintained in years
- No owner after the pilot. An innovation team proves it works and hands it to an operations team with no budget line, no headcount and other priorities
None of those are solved by a better model. All three are organisational, which is why swapping to a newer model rarely rescues a stalled project. The rush to ship AI agents inside every business app runs into exactly the same wall.
The cost overrun problem
The second finding is harder to explain away. Writer's enterprise adoption research found that 79% of enterprises experienced AI cost overruns in the past 12 months.
Pilots are cheap and misleading. A proof of concept runs on a small volume of requests, often on promotional pricing, with engineers who are already on payroll. The per-request cost looks trivial.
Production changes every term in that calculation. Volume multiplies, each request may trigger several model calls rather than one, retries and safety checks add more, and somebody has to be paid to keep it running.
A worked example of the gap
Take a support summarisation tool. In the pilot it handles 500 tickets a month at roughly four cents of model cost each. That is $20 a month. Nobody bothers to forecast it.
Production is 120,000 tickets a month. Each one now makes three model calls instead of one, because the team added a retrieval step and a quality check after early errors. Cost per ticket lands nearer twelve cents.
That is $14,400 a month, or about $173,000 a year, against a pilot that suggested $240. Add an engineer at a fraction of their time to maintain it and the annual figure moves again. The tool can still be worth it, easily, if it saves enough agent time. The problem is that nobody approved a $173,000 line item, so it arrives as a surprise and the project gets paused while finance works out what happened.
If you are building the business case, work the return on investment from production volumes, not pilot volumes. It is the single most common modelling error in this whole category.
What the small minority do differently
The companies getting EBIT impact tend to share a few habits, and none of them are exotic.
They pick a process, not a capability. The failing brief is we should use AI in customer service. The working brief is we want first-response time on billing queries under two minutes without adding headcount. The second one has a number attached, so you can tell whether it worked.
They also build inside the existing tool rather than beside it, so nobody changes what they do. They instrument a baseline before launch, because without one you cannot prove any improvement and the project dies at the first budget review. And they assign a single owner with authority over the workflow itself, not just over the software.
One more pattern: they start where errors are cheap. Drafting, summarising, classifying and routing all have a human checking the output, so a wrong answer costs a moment. Systems that act without review are a much harder second step, and the question of who is responsible when an agent gets it wrong has no settled answer yet.
The spending continues regardless
You might expect a 95% disappointment rate to slow investment. It has not.
Gartner forecasts worldwide AI spending of $2.59 trillion this year, a 47% increase. Agent software specifically is projected at $206.5 billion, rising toward $376.3 billion the following year. Around 59% of companies are putting more than $1 million a year into AI.
Some of that is rational. When a technology is genuinely transformative for a minority of use cases, paying to find out which ones apply to you is defensible. What is not defensible is running the same pilot shape repeatedly and expecting a different ending. The failure modes are known, documented and mostly organisational, which means they are fixable by people who are not machine learning specialists.
Frequently Asked Questions
Why do so many AI pilots fail?
Mostly for non-technical reasons. The tool sits outside the workflow people actually use, the production data is messier than the pilot sample, or nobody owns it once the innovation team moves on. Model quality is usually the part that performed well, which is why the pilot got approved in the first place.
How much should a company budget for an AI project?
Base it on production volume rather than pilot volume, and assume more than one model call per user action once retrieval and quality checks are added. Then add ongoing engineering time for maintenance. The 79% overrun rate comes almost entirely from skipping those three adjustments.
Is it better to build AI tools in-house or buy them?
Buy where the process is standard across companies, such as meeting notes or code completion, because a vendor amortises the work across many customers. Build where the process is genuinely specific to you and is a competitive advantage. Building something a vendor already sells well is the most expensive path to an average outcome.
What is a realistic timeline to see returns from AI?
For a narrowly scoped process with a measured baseline, a few months is reasonable. For anything requiring a workflow change across several teams, plan in quarters and expect the organisational work to take longer than the technical work. Projects promising transformation within weeks are describing a demo.
Fix the process, not the model
The failure rate around enterprise AI is not evidence that the technology is empty. A small group of companies is extracting real earnings from it, and the spending forecasts suggest that group will grow.
What the data does say is that the bottleneck moved. It is no longer the model. It is whether anyone rebuilt a process around it, measured a baseline first, and budgeted for the volume they would actually run. Those are management problems, and management problems are the kind you can fix without hiring a research team.
Written by
Quick Trend Insights Editorial Team
Our editors track the latest in technology, business, finance, and culture, turning fast-moving news into clear, reliable insight you can act on.



