Guide

How to measure AI ROI

Most AI deployments are judged on a feeling. That is how a working system gets cancelled and a useless one gets expanded. The fix is unglamorous: pick two or three numbers you already track, write them down before you start, and look again at ninety days.

Write down the before

This is the whole method and almost nobody does it. Before anything is deployed, record what the workflow costs today in one line. Not a spreadsheet — a sentence.

"We get roughly 40 enquiries a week, about a third arrive outside office hours, median first reply is 6 hours, and roughly 5 a week go unanswered entirely."

That sentence is worth more than any dashboard, because in three months it is the only thing standing between you and an argument about impressions. Take ten minutes.

Three numbers that are hard to argue with

  • Response time on inbound, including out of hours. Directly attributable, immediately measurable, and for most businesses the one with real revenue attached. If AI does not move this, it is not doing the job.
  • Hours your most expensive person spends on work that does not need them. Ask them, honestly, for a fortnight before and a fortnight after. Imprecise and still more useful than anything automated.
  • Things that used to fall through. Unanswered enquiries, uncontacted leads, reports nobody wrote, follow-ups nobody made. Often the largest real effect, and the one businesses forget to count because the baseline was invisible.

Two or three is the right number. A measurement framework with twelve metrics gets abandoned in week three.

The metrics to ignore

  • Messages handled. Volume is not value. A system handling a thousand messages badly scores well.
  • Tokens or model usage. That is a cost line, not a return.
  • "Time saved" as a vendor-reported figure. Usually messages multiplied by an assumed handling time the vendor chose.
  • Accuracy percentages with no denominator. Accurate against what, measured by whom?
  • Deflection rate. Popular and dangerous — it rewards not reaching a human, which is not the same as the customer being helped.

Counting the costs properly

Three, and only the first is on the invoice: the platform, the model usage (consumption-based, which is why hard caps matter more than a low headline), and attention — choosing the process, owning the integration, reviewing output early, adjusting as the business changes. The third is the largest and the one nobody budgets.

Do count the review tax honestly in the first month, and do note that it declines. A business case that assumes zero review is wrong; one that assumes the week-one level forever is also wrong.

The ninety-day review

One meeting, four questions:

  1. Did the before-number move? Against what you wrote down, not against memory.
  2. What is it doing that we did not expect? Often the most valuable finding, in both directions.
  3. What are we still correcting? Every repeated correction is a procedure that has not been written down yet.
  4. Would we buy it again at full price? The cleanest possible test, and the one everything else is a proxy for.

What a negative result looks like

Worth defining in advance, because a deployment nobody is willing to call a failure runs forever. Three signals, any of which is enough:

  • The before-number has not moved and nobody can say why.
  • The review burden has not declined since week two — meaning corrections are not being captured as procedures, so it will never compound.
  • People have quietly gone back to the old way. This one is decisive, and you will only notice it if you ask.

A negative result on one workflow is not a verdict on the technology. It usually means the workflow was the wrong choice, which picking a better one fixes.

Related reading

FAQ

How do you measure ROI on an AI deployment?

Write down what the workflow costs today in one sentence before you deploy, pick two or three numbers you already track, and compare at ninety days. The three hardest to argue with are response time on inbound including out of hours, hours your most expensive person spends on work that does not need them, and things that used to fall through entirely.

Which AI metrics are misleading?

Messages handled, since volume is not value. Token or model usage, which is a cost rather than a return. Vendor-reported time saved, usually volume multiplied by an assumed handling time. Accuracy percentages with no denominator. And deflection rate, which rewards not reaching a human rather than the customer being helped.

What costs should I include?

Platform, model usage, and attention. The third — choosing the process, owning the integration, reviewing output early, adjusting as the business changes — is the largest and the one nobody budgets. Count the review tax honestly in month one and note that it declines.

How do I know if an AI deployment has failed?

Three signals, any one of which is enough: the before-number has not moved and nobody can say why; the review burden has not declined since week two, meaning corrections are not being captured so it will never compound; or people have quietly gone back to the old way. The last is decisive and you will only find it by asking.

Agree the number before the build

We scope the workflow and the measure together, so there is something to judge the result against that neither of us invented afterwards.

Start a deployment