Measuring AI when ROI doesn't fit
Payback-period ROI suits projects that replace a known process. A lot of AI isn't that. Here is what I'd track alongside it, in terms finance can audit.
Published Updated 4 min read
The standard way to justify a technology investment is simple: estimate the savings, divide by the cost, check the payback period. It works well when you're replacing a process you already understand. Applied to AI, it tends to go wrong in one of two directions. Either good projects get killed because year one doesn't pay back, or people inflate the numbers until it does.
Plenty of organizations are stuck here. In a 2024 Gartner survey, 49% of respondents named difficulty estimating and demonstrating value as the top barrier to adopting AI. A year later, S&P Global found that the share of companies abandoning most of their AI initiatives had jumped to 42%, from 17%. The companies in that survey blamed cost, privacy and security, not measurement. Still, I suspect some of those projects were judged on a yardstick that couldn't see what they were for.
Why payback math misses so much
AI is partly automation and partly platform. The automation part, where a task gets done faster or cheaper, fits a normal business case. The platform part, where a new capability makes other things possible, mostly doesn't. Four things get lost:
- Value nobody planned for. A support assistant built to deflect tickets turns out to be the best source of product feedback you have. The business case never counted that, so nobody gets credit for it.
- Compounding. A tool can lose money in its first year while the people using it get steadily better with it and other teams copy the approach. A point-in-time calculation can't see a curve.
- Tangled cause and effect. The gains usually come from AI plus a redesigned process. Trying to isolate the AI's share is like valuing an engine separately from the transmission.
- The wrong horizon. A 12-month payback rule applied to a capability that takes two or three years to build will reject it every time.
Finance and technology are using different rulers
In EY's 2025 technology risk poll, 56% of CFOs said AI integration was a top priority for the next two to four years, against 70 to 72% of CIOs and CTOs. My read is that the two groups are measuring different things. Finance asks what it costs per user, when it pays back and whether the decision can be reversed. Technology asks what it makes possible and what it costs not to have it.
The expensive part is what happens in the gap. When approval feels like "no" dressed up as "not yet", departments buy their own tools. The CFO thinks spending is contained because the IT budget is flat. The CIO thinks risk is contained because formal approvals are required. Often neither is true, because the real spending and the real risk sit in budget lines nobody is watching. That is how shadow AI starts.
What I'd measure alongside ROI
I'd keep the business case and add three measures that capture what it misses.
The first is learning speed: how long it takes to go from an idea to something running, how many experiments are active, and how often one team's solution gets picked up by another. It is the one that moves first.
The second is decision coverage, meaning the share of work that gets expert-level attention. If senior staff used to review a tenth of customer conversations, and an AI-assisted process lets them check all of them for the same patterns, that is a change you can count.
The third is compounding: new use cases built on earlier ones, how quickly new hires get up to speed, and how much gets reused instead of rebuilt.
Each of these should come with a translation into money, even a rough one: months of ramp-up saved times what a new hire costs during ramp-up, or errors avoided times what an error costs. The translation is what lets finance audit the claim instead of taking it on faith.
It also helps to watch leading and lagging measures together. Leading indicators without lagging ones are hope. Lagging indicators without leading ones only tell you what already happened. If people are using their AI budgets heavily but productivity isn't moving, you have an execution problem. If productivity is up but hardly anyone is experimenting, the value is concentrated in a few use cases and won't spread. Budget use is a surprisingly good signal on its own: well under half usually means too much friction or too big a budget, and close to 100% means either strong demand or too small a budget. My rule of thumb is to aim for something like 70 to 85%.
Start with a spreadsheet
None of this needs new software. A form to log experiments, a shared table that tracks each one from idea to production, a monthly note on what worked and a quarterly roll-up in dollar terms will do for the first six months. If you already keep a central record of AI work to avoid building the same thing twice, most of the data is there.
Then test the approach before arguing about it. Pick five to ten current initiatives and assess them both ways for a quarter: the usual ROI calculation and the measures above. Where the two disagree is where the interesting conversations are.
One exercise I'd run first, because it takes an hour: take a single AI implementation that everyone agrees worked. Ask the finance lead and the technology lead, separately, two questions. What made it financially viable? What capability did it build? If their answers barely overlap, you've found the gap, and it probably explains why other projects are stuck in pilot purgatory.
What I don't have is a clean way to price optionality, the value of things you can now do but haven't done yet. For now I'd rather put a rough number on it and argue about the number than leave it out and let the payback math decide by default.