The first AI system an operator should build is not a coding assistant, and it is not a general chatbot pointed at the whole company. It is a narrow tool that automates one specific, repetitive task the team already does by hand, several times a week. OpenAI's own enterprise adoption guide groups nearly every generative AI deployment into six buckets, content creation, research, coding, data analysis, ideation and strategy, and automation, and it points new adopters toward automating tedious, repetitive work first, things like answering similar emails or sorting support tickets, rather than starting with the flashiest use case (OpenAI).
Skip the coding agent
Most operators default to a coding agent as their first build, because that is where the AI conversation has been loudest. The data says that instinct is backwards for almost everyone. Citing Anthropic's own usage numbers, Y Combinator president Garry Tan noted that software engineering accounts for 49.7% of all agentic AI tool calls, while every other vertical, healthcare, legal, sales operations, sits under roughly 5 to 9% (Garry Tan). That is not evidence coding agents fail. It is evidence the category is already saturated with builders, while the actual repetitive bottleneck inside a given company, the ticket queue, the intake form, the weekly report, usually sits in one of the underserved 5 to 9% buckets, untouched.
Bounded beats broad
The reason a narrow tool outperforms a general assistant is not motivation, it is measurement. In a controlled study of 95 professional developers, those given GitHub Copilot on a single bounded task, building an HTTP server in JavaScript, finished 55% faster on average, 1 hour 11 minutes versus 2 hours 41 minutes, a result strong enough to be statistically significant at p=.0017 (GitHub). The task had a fixed input and a checkable output. That is what makes the time saved real and countable.
Broad assistants do not get that clean a result, because the input keeps changing shape. Glean's Work AI Index, surveying 6,000 full time digital workers across the US, UK, and Australia, found employees report AI saves them close to 11 hours a week, but they lose an average of 6.4 hours a week botsitting, feeding the tool context, watching its output, and fixing what it got wrong (Glean). Net the two numbers against each other and the real gain falls under 5 hours, close to what a well scoped tool delivers on its own, with none of the supervision overhead.
What it looks like at scale
McKinsey's internal platform, Lilli, is the clearest example of scoping this correctly. Rolled out firmwide in July 2023, Lilli was not built to do everything, it was built for search and synthesis work specifically. Consultants use it roughly 17 times a week each, report up to 30% time savings on that task, and the firm estimates it recovers about 50,000 labor hours every month, with 72% of the firm active and generating more than 500,000 prompts monthly (McKinsey). One job, done well, adopted by almost everyone. That is the model to copy, not a platform that tries to be a coworker for every task at once.
Decide where the hours go before you build
None of this matters if the freed time disappears into the same low value work it replaced. Gartner surveyed 210 chief sales officers and senior sales leaders between January and February 2026 and found AI saves sellers an average of 4.8 hours a week, yet 72% of sales organizations never reinvest that recovered time into higher value activity (Gartner). The tool worked. The organization never had a plan for the hours it created.
This week, before building anything: name the one bounded, repetitive task your team does by hand most often, the same three email replies, the recurring report pull, the ticket triage, and write down, in advance, exactly what the recovered hours will be redirected to. Build the narrow tool for that single task first. Skip the coding agent and skip the do everything assistant. Scope it like Lilli, not like a general chatbot, and hold the team to the plan for the hours it frees, or they will vanish the way Gartner found they do for most sales teams.