The AI readiness assessment passed. Tool deployment is above target, workforce training has cleared its milestones, and the executive committee received a green paper with confidence. Eighteen months of pilots have run cleanly. The productivity gains reported in the pilot summaries have not shown up in the enterprise numbers, and the CFO is now asking why the case that was made for the investment is not visible in the case being reported back.
This is the pattern most enterprises are living inside right now. BCG’s 2025 analysis of AI value capture found that only 5% of enterprises, the segment they call “future-built”, are converting AI investment into 1.7x the revenue growth and 3.6x the total shareholder return of everyone else (Apotheker et al., 2025). The other 95% sit across three segments (scaling, emerging, and stagnating) that have deployed AI, trained their people, and are not moving the enterprise number. McKinsey’s most recent state-of-AI research finds the same shape: around 80% of organisations have deployed generative AI, and only 5.5% to 6% attribute more than 5% of EBIT to it (McKinsey, 2025). Adoption is doing what the executive team asked of it. The financial return sits somewhere else in the system.
The scorecard measures the wrong layer
Readiness assessments answer a specific question well: can the organisation adopt AI. They test the presence of tools, the coverage of training programs, the existence of policy documents, and the volume of pilots underway. Every one of those signals matters, because they establish that the organisation has cleared the threshold at which AI use is possible. A green scorecard confirms that threshold has been cleared.
The distinction that matters for returns is adoption breadth versus integration depth: adoption breadth measures how widely AI tools have been rolled out across the workforce, while integration depth measures how deeply investment governance, accountability, and decision rights have been restructured around what AI actually changes. Readiness assessments measure the first. Returns come from the second.
What the same assessment does not test is whether the organisation can convert adoption into returns. The mechanisms that produce financial value from AI are structural: how investment decisions are made and revised, where accountability for AI-enabled outcomes sits, which decision rights are being redistributed as agents take on work, and how the underlying processes are being redesigned around what AI actually changes. Deloitte’s 2026 State of AI in the Enterprise research surveyed 3,235 leaders across 24 countries and found that only around one-third of organisations are doing the deep process redesign that turns AI capability into value, while roughly the same proportion is using AI at a surface level with little or no change to how work is organised (Deloitte, 2026). The scorecard cannot see that distinction, because it was not commissioned to.
The DORA 2025 State of AI-assisted Software Development report describes the mechanism precisely: AI is an amplifier, and it magnifies both the strengths of high-performing organisations and the dysfunctions of struggling ones (DORA, 2025). An organisation with clear investment governance and defined accountability for outcomes uses AI to accelerate work that was already coherent. An organisation where accountability is diffuse and decision rights are unclear uses AI to produce more of the same unresolved work, faster. Both organisations look identical on a readiness scorecard. Only the first produces returns.
The tools miss integration depth
The layer the scorecard cannot see is the layer the scorecard was not commissioned to find. Readiness assessment tooling is largely built by vendors whose commercial interest is adoption, so the tools measure what the vendor’s product can move. Integration depth is harder to score, harder to standardise into a maturity model, and harder to sell as a diagnostic product. It is also the layer the vendor cannot fix on the client’s behalf, which weakens the commercial logic for measuring it in the first place.
Inside the organisation, the same pattern reinforces itself. Adoption metrics are visible, standardised, and can be reported quarterly. Integration depth requires a different kind of evidence: what has the investment committee approved and revised in the last two cycles; how many AI-enabled initiatives have a single, accountable owner rather than a cross-functional working group; how many core processes have been redesigned around what AI changes rather than layered on top of unchanged process. That evidence is harder to produce and slower to move. Given the choice between a metric that ticks up cleanly and a metric that requires structural work to shift, most quarterly reporting settles on the first.
A passing scorecard does not predict returns
That reporting bias is why a passing scorecard tells the executive committee less than it appears to. A pass tells the committee that the organisation has the tooling and the trained workforce to run AI pilots. That is a real threshold and worth clearing. What it does not tell the committee is whether the organisation has the investment governance to allocate capital across a portfolio of AI initiatives with revision, whether accountability for each initiative’s outcome sits with a single owner rather than a working group, or whether the decision rights structure has actually shifted to determine which decisions AI-enabled agents will make and which decisions humans will keep. Those conditions are where returns come from, and they are the conditions the future-built segment has built while the rest of the market has not.
The next investment cycle is where the exposure becomes expensive. When the CFO returns for the second-round approvals, the case for continued AI investment rests on what the first round produced. If the enterprise numbers have not moved, the argument has to bridge from “we passed the readiness assessment and ran the pilots” to “and here is why the returns have not appeared yet”. That bridge is where the CIO becomes exposed, because the readiness score anchored the committee’s expectation to a threshold that does not predict returns. The distance between the reported readiness and the reported value widens with every quarter the structural condition stays unaddressed, and by the time the board asks about it, the answer has to explain both the shortfall and why it was not visible earlier.
A second opinion tests structural capacity
Because that exposure begins at the next approval round, the diagnostic the executive committee actually needs before it opens is not another readiness score. What the committee needs is an assessment of whether the organisation has the structural capacity to generate returns from what it is already doing.
A useful second opinion examines investment governance in operation: what the committee has approved and revised across the last two cycles, and how quickly capital moves when a pilot underperforms. It goes to accountability, asking whether AI-enabled outcomes have a single owner who can be held to a result, or whether they sit inside a working group that cannot. And it goes to decision rights, which are the hardest layer to see from the inside, because it tests what has actually shifted about who decides what, now that agents can execute work that used to require human judgment. Framework names matter less here than what is being tested. McKinsey’s research identifies the organisational conditions that separate value-capturing companies from adopters: 55% of high performers have reworked core processes, against roughly 19% of everyone else (McKinsey, 2025). That is the layer a second opinion has to reach.
The readiness assessment did the job it was built to do. It confirmed the organisation can run AI pilots. The reason the enterprise numbers have not moved is that the assessment did not measure the layer where returns actually appear. A second-opinion diagnostic scoped to investment governance, accountability, and decision rights, initiated before the next approval round, gives the executive committee the evidence the readiness score cannot provide, and gives the CIO the bridge from adoption to returns before the CFO’s question hardens into a board question. The executive who sees this distinction before the next cycle is working with a different investment case than the one who inherits the readiness score as the whole answer.
What this means for senior leaders
- Treat a passing readiness assessment as confirmation of adoption capacity, not as evidence of return-generating capacity. The two are measured differently, and only the second predicts financial outcomes.
- Before the next AI investment approval round, commission a diagnostic scoped explicitly to investment governance, single-owner accountability, and decision rights redistribution. That is the layer where the future-built 5% differ from everyone else.
- Assume the readiness score has anchored the executive committee’s expectation to the wrong threshold. The bridge from adoption to returns has to be built in the committee’s language before the CFO builds it in the board’s.
- Use the DORA amplifier finding as a discipline: if accountability was diffuse before AI, it will be more diffuse after. Fix the operating model condition first, then let AI amplify a coherent structure rather than an unresolved one.
- Report integration depth alongside adoption metrics from this quarter onward. If the metric that predicts returns is not on the executive dashboard, the executive team cannot manage to it.
References
- Apotheker, S., et al. (2025). The widening AI value gap. Boston Consulting Group.
- Deloitte. (2026). State of AI in the enterprise (8th ed.). Deloitte.
- DORA. (2025). State of AI-assisted software development. DORA.
- McKinsey. (2025). The state of AI. McKinsey & Company.