In February 2024, Klarna announced its OpenAI-powered customer service assistant had handled 2.3 million conversations in its first month live, doing the work of 700 full-time agents across 23 markets and more than 35 languages, and cutting average resolution time from 11 minutes to under 2, according to a Klarna press release. Fifteen months later, CEO Sebastian Siemiatkowski reversed course, publicly acknowledging the AI-driven push had hurt service quality, and started rehiring human agents after customers pushed back on complex interactions, per Forbes.
Same company. Same model generation. Same underlying technology. What changed wasn't the AI. It was the scope of what got automated versus what should have stayed human, and nobody caught the mismatch until customers did.
That gap, between a pilot that looks like a success on a dashboard and one that survives contact with the full range of real customer problems, is the actual story behind the number everyone quotes. MIT's NANDA initiative found that 95% of generative AI pilots at companies deliver no measurable ROI or P&L impact, despite an estimated $30 to $40 billion in enterprise generative AI spending, per Fortune's coverage of the report. Read as proof that AI doesn't work, that number is misleading. Read as evidence that most companies are automating the wrong slice of a process, it's a diagnosis with a fix.
Buy the pain point, don't build the platform
The same MIT NANDA report found something operators tend to skip past: companies that bought AI tools from external vendors and partnered succeeded about 67% of the time, roughly twice the success rate of internally built tools, according to Fortune. Lead author Aditya Challapally put it plainly: companies that succeed "pick one pain point, execute well, and partner smartly with companies who use their tools."
That's the opposite of what most executive teams reach for. Building internally feels like it protects IP and avoids vendor lock-in. In practice it means your team is debugging prompt engineering and integration edge cases instead of buying that expertise from a vendor who already shipped it to a hundred other customers. Gartner predicted at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs, and unclear business value, per the Gartner newsroom. Every one of those four causes is a scoping and process failure, not a model failure.
The 70/20/10 split nobody budgets for
BCG's research on AI value creation found that only 26% of companies have built the capabilities to move past proof of concept, and just 4% report substantial value creation. The same research attributed roughly 70% of scaling failures to people and process issues, 20% to technology, and only 10% to the AI algorithms themselves, according to BCG. Most AI budgets are allocated in the opposite ratio: heavy spend on the model and the integration, almost nothing on redesigning the workflow around it or training the team that has to catch what it gets wrong.
Klarna's assistant handled the 2.3 million easy, repeatable conversations well. What it didn't have was a tested process for routing the genuinely hard cases back to a human before the customer got frustrated enough to complain publicly. That's a process gap, and BCG's data says process gaps are where 70% of these projects actually die.
McKinsey's most recent State of AI survey, fielded in the summer of 2025 across nearly 2,000 respondents in 105 countries, found 88% of organizations report regular AI use in at least one business function, but only 6% say AI has produced significant enterprise-wide EBIT impact, and nearly two-thirds haven't started scaling AI across the enterprise at all, per McKinsey. Usage and value are not the same metric, and most companies are only measuring the first one.
"Executives are impatient to see returns on GenAI investments, yet organizations are struggling to prove and realize value," Gartner VP Analyst Rita Sallam said in the firm's July 2024 release on project abandonment.
What to do this week
Pull up every AI pilot currently running in your company and check it against three questions the data above actually supports. Is it scoped to one specific, measurable pain point, or is it a platform play dressed up as a pilot? Are you buying that capability from a vendor who has already solved it elsewhere, or paying your own team to relearn it? And is there a defined, tested path for routing the cases the AI can't handle back to a human, before a customer finds that gap for you?
If you can't answer all three, you don't have a pilot. You have a demo running in production, and the 95% failure rate says that's exactly what's failing.
- Scope to one pain point before adding a second use case.
- Default to buying from a vendor with existing customers over building in house.
- Build and test the human escalation path before launch, not after complaints.
- Measure P&L impact, not usage rate, as the success metric.