Klarna's AI customer service assistant handled 2.3 million conversations in its first 30 days live, doing the equivalent work of 700 full-time agents and automating two-thirds of all customer service chats, according to Klarna's own announcement in February 2024. Fifteen months later, CEO Sebastian Siemiatkowski was hiring humans back, telling Forbes in May 2025 that the AI-first push had produced "lower quality" service and that customers wanted to talk to people. Same company, same technology, two opposite verdicts less than two years apart. Founders trying to figure out how to actually use AI inside their business need to understand why that swing happened, because most of them are about to repeat it.
The data says Klarna's whiplash is the norm, not the exception. MIT NANDA's "The GenAI Divide: State of AI in Business 2025," built on 300 public AI deployments, 52 structured interviews, and 153 leader surveys, found that despite $30 to $40 billion in enterprise generative AI investment, 95% of organizations got zero measurable return. Just 5% of integrated pilots extracted millions in value, per Fortune's August 2025 coverage of the report. McKinsey's global survey of 1,993 executives across 105 countries, fielded between late June and late July 2025, found 88% of organizations now report regular AI use in at least one business function, up from 78% a year earlier, yet only 39% can point to any measurable EBIT impact, according to McKinsey's State of AI report. PwC's 29th Global CEO Survey of 4,454 CEOs across 95 countries, fielded September through November 2025 and released in January 2026, put a number on the gap between adoption and payoff: 56% of CEOs report no revenue or cost benefit from AI at all, while only 12%, the "vanguard," achieved both lower costs and higher revenue, per the PwC report.
The variable that actually explains the split
MIT NANDA's report isolates the mechanism: internally built generative AI systems succeed roughly one-third as often as tools purchased from specialized vendors, meaning in-house builds fail about twice as often as externally sourced deployments, per the MIT NANDA study. That is the Klarna pattern in miniature. Klarna didn't buy a narrow, vendor-built tool for one workflow. It built and ran a custom system aimed at replacing an entire function, and the automation numbers looked spectacular right up until quality complaints forced a partial reversal. Ambition scaled faster than the tooling could support it.
Buy the narrow tool, don't build the department
The founder-level lesson is not to avoid AI. It is to match the build decision to the stakes. A ten-person startup does not have the QA infrastructure, the eval pipelines, or the headcount to build and monitor a custom AI system the way a platform team would. Buy a specialized vendor tool for one well-defined, high-frequency process: support triage, lead scoring, contract review, expense coding. Hold off on building anything custom until that narrow deployment has a track record, then expand only after the metric proves out. Building in-house should be reserved for the parts of the business that are the actual product, not the back office running around it.
Measure before you scale, not after
Gallup's Q4 2025 workplace data shows what happens without that discipline: daily AI use among U.S. employees crept from 10% to 12%, even as overall usage growth flattened, per Gallup. Adoption is stalling at low, un-scaled levels because most rollouts were never tied to a metric that would justify pushing further. The founders in PwC's 12% vanguard and McKinsey's 39% EBIT-impact group share one habit: they picked the number before they picked the tool, cost per resolved ticket, hours saved per contract review, conversion lift per qualified lead, and only expanded the deployment once that number moved.
Where the exception holds: AI as the founding team's core infrastructure
There is one place where founders should build rather than buy: the product itself, from day one. A quarter of startups in Y Combinator's Winter 2025 batch reported that 95% of their codebase was AI-written, and YC CEO Garry Tan noted companies reaching $10 million in revenue with teams of fewer than 10 people, according to TechCrunch. That is not a contradiction of the build-versus-buy finding above. Those founders are the vendor. AI is core infrastructure they own and iterate on directly, not a layer bolted onto an existing process by a team without the context to fix it when it breaks.
Match the build decision to the stakes: buy the narrow tool for the back office, build only what is the product itself.
The action for this week: list every AI initiative currently running in the business. For each one, write down the metric it is supposed to move and whether that metric has actually moved. Anything without a metric gets killed or narrowed to a single process by Friday. Anything with a metric that is working gets the budget the un-measured pilots were quietly consuming.