Article

You're Measuring AI Completely Wrong

Your AI metrics are usage stats, output counts, and adoption rates. None of them tell you whether AI is making your operation better. The only metrics that matter are operational ones - and almost nobody is tracking them.
A carnival funhouse hall of mirrors with a single object reflected in multiple distorting mirrors, each showing a different warped version, under garish coloured lighting

Stuart Totterdell

Technical Director

How is your AI performing?

If your answer involves any of the following, you are measuring it wrong: number of users. Number of queries processed. Volume of outputs generated. Adoption rate. User satisfaction score. Time spent in the platform.

These are activity metrics. They tell you that AI is being used. They tell you nothing about whether it is making your operation better.

And yet they are the metrics that almost every mid-market business reports when asked about AI enablement performance. Because they are the metrics the vendors provide, the dashboards display, and the internal champions present to the board.

They are also completely useless for determining whether your AI investment is delivering value.

The vanity metric problem

Activity metrics feel meaningful because they are always positive. Usage goes up. Queries increase. Outputs accumulate. The charts trend upward. The vendor sends a congratulatory email about adoption milestones.

But activity is not value. A customer service team using an AI chatbot fifty times a day is activity. Whether those fifty interactions resolved issues faster, reduced escalations, or improved client retention - that is value. And those two things are measured completely differently.

Activity metrics tell you that the tool is being used. They do not tell you that the tool is useful.

The distinction matters because activity can increase while operational performance stays flat - or even gets worse. A team might use an AI summarisation tool extensively and still take the same time to process cases, because the summaries are not connected to the workflow and still require manual action. Usage is high. Impact is zero.

What you should be measuring

The only metrics that matter for AI are operational metrics - the same metrics you would use to measure any process improvement:

Time. How long does the process take now versus before AI was applied? Not the individual AI task - the entire process from trigger to completion. If AI summarises a document in ten seconds but the person still takes forty minutes to act on it, the AI has not changed the process time.

Errors. Has the error rate in the process changed since AI was introduced? If AI classifies incoming requests, how many does it classify correctly? How many require human correction? If the correction rate is twenty percent, that is not an AI success - it is a quality problem that happens to be powered by AI.

Cost. Has the cost per transaction, per case, per order changed? Not the cost of the AI tool itself - the total operational cost of the process it is embedded in. If the AI costs £500 per month but saves £2,000 in labour and error correction, it is delivering value. If it costs £500 per month and the process costs the same as before, it is not.

Cycle time. How quickly does work move through the process? If AI was supposed to accelerate decision-making, are decisions actually being made faster? Measured from when the need arises to when the decision is implemented - not from when the AI generates its output.

Capacity. Has the team's capacity to handle volume increased? If AI was deployed inside a business automation workflow to scale a process, can the team now handle more work without more people? Or has the workload remained the same, with AI simply adding a new step to an unchanged process?

Why nobody tracks these

There are three reasons operational metrics are rarely tracked for AI projects.

The first is that nobody measured the baseline. The AI was deployed without first quantifying how the process performed without it. There is no "before" to compare against. So the team defaults to the metrics they can easily access - the ones the AI platform provides - which are all activity metrics.

The second is that the operational metrics are harder to isolate. Process improvements are influenced by many factors - staffing changes, process redesign, seasonal variation. Attributing a specific improvement to AI requires a controlled comparison that most businesses do not set up. It is easier to say "AI usage is up thirty percent" than to prove "AI reduced our classification error rate from eight percent to two percent."

The third is that operational metrics might reveal uncomfortable truths. If the AI has been deployed for six months and the process time has not changed, the error rate has not improved, and the cost per transaction is flat, the investment has not delivered value. That is a difficult conversation - especially if the AI was championed by a senior leader or the vendor relationship is ongoing.

So the business reports activity metrics, the charts look positive, and nobody asks whether anything in the operation has actually improved.

The measurement framework

Before deploying any AI, establish three to five operational metrics for the process it will affect. Measure them for at least a month before AI is introduced. Document the baseline.

After deployment, measure the same metrics. Monthly. For at least six months. Compare them directly to the baseline. Be specific: not "things feel faster" but "average processing time reduced from 4.2 hours to 1.8 hours."

If the metrics have not improved after three months, the AI is not delivering operational value - regardless of how much it is being used. Either the deployment needs to be redesigned, the underlying IT and process strategy needs to be fixed first, or the AI is not the right tool for the problem.

Report operational metrics alongside activity metrics. If usage is high and operational improvement is zero, that is important information. If usage is moderate but operational improvement is significant, that is a success story - even if the adoption dashboard does not look as impressive.

The real test

The real test for any AI investment is simple: has it changed an operational metric that matters to the business?

Not "is it being used." Not "do people like it." Not "does the vendor report high engagement."

Has it reduced the time a process takes. Has it reduced the errors. Has it reduced the cost. Has it increased the capacity.

If you cannot answer those questions with specific numbers, you are not measuring AI. You are measuring activity. And activity, on its own, is worth nothing.

A carnival funhouse hall of mirrors with a single object reflected in multiple distorting mirrors, each showing a different warped version, under garish coloured lighting

Can you prove your AI investment is delivering operational value?

We help you set the right baselines, measure what matters, and deploy AI that changes process metrics - not just usage charts.

Measure what matters