Your AI pilot worked. The board approved scaling. Six months later, the roll-out has not landed. Usage is patchy, the productivity numbers from the pilot are not reproducing at scale, and the team is quietly asking whether the tool was oversold. The instinctive move at this point is to approve another pilot, this time closer to the point of value, to rebuild confidence.
In Brief
- AI pilots stall at scale because the AI operating model they need to land in — the roles, workflows, decision rights, and performance measures — was never redesigned.
- Approving a second or third pilot answers a question the first already answered, and defers the operating model decisions blocking production.
- McKinsey identifies workflow redesign as the largest EBIT lever from generative AI; Forrester puts 40% of organisations in POC purgatory.
- The next executive act is initiating the AI operating model work for pilots already in the gap, not approving another proof point.
That instinct is what keeps organisations in this position for another eighteen months.
The pilot did not fail. Neither did the next one it justified. What did not happen, and has still not happened, is the redesign of the AI operating model the pilot needs to scale into: the roles people are hired into, the workflows they run, the decisions they are allowed to make on their own, and the incentives their line manager is measured against. Those four elements are what the sales conversation calls the operating model as if it were a single lever. It is not one lever. Each element is a separate capital decision an executive is being asked to make in sequence, and the pilot was never designed to make any of them for the rest of the organisation. It was designed to prove the technology works. It proved that. Pilots are good at that job.
McKinsey (2026) found that 88% of organisations are now experimenting with AI while 81% report no meaningful bottom-line impact. The same McKinsey research identifies workflow redesign as the single largest EBIT lever from generative AI, and finds that high performers are approximately three times more likely than others to have significantly modified workflows around the technology. Forrester (2026) puts 15% of organisations in territory where AI agent deployments are producing measurable return, and 40% in what Forrester analysts have taken to calling POC purgatory: a state where the pilots keep running but the organisation cannot get past them. Neither firm blames the technology. Both point at the receiving organisation.
The reason the receiving organisation does not get redesigned is that pilots are structured to make redesign optional. A pilot is scoped narrowly, to one team, one workflow, one measurable outcome, precisely so that the surrounding organisation does not have to change to accommodate it. That scoping is what makes a pilot cheap enough to approve and fast enough to prove. It is also what makes its success a poor guide to what production requires. Success at pilot scale demonstrates that a capable person, working alongside the technology inside a bounded context, can produce a better outcome than the baseline. Production is a different exercise entirely. It requires that hundreds of less experienced operators, working across contexts designed for a pre-AI logic of work, produce that outcome as a matter of routine. The distance between those two states only becomes visible when the executive tries to close it.
Approving more pilots widens the operating model gap
The pattern that surfaces from here plays out slowly. Because the pilot proved the technology, the sponsoring executive receives a validated business case and permission to scale. Because scaling requires redesigning the workflow the technology is landing in, and redesigning the workflow means renegotiating who decides what and revising the numbers a manager is measured on, the executive who tries to move directly to production runs into resistance from every function whose accountability is about to shift. Legal has not agreed to the new liability model. Finance has not agreed to the new cost allocation, and HR has not signed off on the changed role profiles. The line manager, whose team is doing the work, has not agreed to the changed performance metrics either. Confronted with resistance on this many fronts at once, the default response is to initiate another pilot in a different business unit, on a slightly different workflow, to build a second proof point.
That is where the eighteen months go. Every additional pilot narrows the technology risk that was already low, and leaves the operating model decisions untouched. Dataiku’s 2026 CEO survey (Harris Poll fieldwork) recorded a drop in CEO confidence in deploying AI agents at scale from 41% to 31% year-on-year, and found that 80% of CEOs believe their own role is at risk by the end of 2026 if their AI strategies fail. The confidence drop is not a statement about the technology. It is a statement about the executive’s ability to redesign the organisation around it, which is a different capability from the one that approved the pilot in the first place.
Four decisions the pilot never made
Four decisions sit at that level and nowhere else. Each is a capital act the executive must take before the pilots already inside the organisation have somewhere to land. Together, these four decisions constitute the AI operating model redesign that pilots cannot make for the organisation.
Decision #1: Which roles are being redesigned
McKinsey’s (2026) work on the agentic organisation identifies approximately 75% of current roles as requiring reshaping and new skill mixes to operate alongside AI agents at production scale. McKinsey names three emerging role profiles that most organisations have not yet defined, funded, or hired: agent orchestrators, hybrid managers, and AI coaches. Until those roles exist, the agent has no accountable human counterpart. When an outcome fails, the organisation cannot tell whether the technology failed or whether the human system around it failed, because the human system around it was never built.
Decision #2: Which workflows are being rebuilt
The pilot’s workflow was built to make the technology visible. The production workflow has to make the technology invisible, embedded in the sequence of work in a way that does not require the operator to switch context to use it. That is skilled work. It is what McKinsey’s research points to when it identifies workflow modification as the primary EBIT lever from generative AI. The technology team cannot do it on its own, and it cannot be run in parallel with another pilot. The workflow team and the pilot team compete for the same operators.
Decision #3: Which decisions the agent may make
Every agentic workflow moves some judgement from a human to a system, and some judgement from an individual to a committee. Until the executive has decided which decisions the agent may make on its own, which it must escalate, and where the escalation lands, the operator’s default will be to escalate everything. That reproduces the pre-AI decision load exactly and delivers none of the promised speed. Decision rights redesign is where the productivity case either lands or gets silently negated.
Decision #4: Which incentives are being reset
Line managers optimise against the numbers they are measured on. When the agent produces the outputs a manager is measured on, adopting it becomes a rational act. When the agent automates the activity a manager is measured on, quietly not using it becomes the rational act instead. Incentive design is where AI production is either supported or silently rejected inside the organisation. It is also the decision the executive is least often asked to make explicitly. It sits with HR and Finance rather than with the sponsor, and it moves last.
Redesign the receiving organisation
The executive whose pilot has worked is not looking at a technology problem. They are looking at four capital decisions their organisation has never made, and the pilot proved neither the case for making them nor the specifics of how. Approving a second pilot answers a question that has already been answered. Initiating the AI operating model work — redesigning the roles, workflows, decision rights, and incentive structures around the pilots already in the gap — answers the question that is actually open.
What this means for senior leaders: The next capital decision is not another pilot approval. It is the decision to fund the redesign of the receiving organisation for the pilots that have already worked, starting with the decisions the agent may make, the incentives that should now measure something different, and the roles that need to exist before the year closes. Executives who make that decision this quarter will be working with an organisation that can absorb AI at scale by the end of the following one. Executives who approve a third pilot instead are running the experiment they have already run twice, and their competitors have already stopped running it.
References
- Dataiku. (2026). Global AI confessions report: CEO edition 2026 (Harris Poll fieldwork). Dataiku. https://pages.dataiku.com/global-ai-confessions-ceo-edition
- Forrester Research. (2026). The state of agentic AI in 2026: Companies are chasing, few are catching. Forrester. https://www.forrester.com/blogs/the-state-of-agentic-ai-in-2026-companies-are-chasing-few-are-catching/
- McKinsey & Company. (2026). The state of AI: How organizations are rewiring to capture value. McKinsey & Company. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- McKinsey & Company. (2026). Six shifts to build the agentic organization of the future. McKinsey & Company. https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/the-organization-blog/six-shifts-to-build-the-agentic-organization-of-the-future