How to measure AI results
An AI rollout can be measured at 3 layers: whether the output is good, whether the work now happens differently, and whether the business moved. Each is harder than the one before, and most rollouts measure none of them, which is why the question of whether an AI tool is working so often gets settled by whoever speaks with most conviction.
AI is genuinely hard to measure. The outputs are fuzzy. The baselines are unclear. The benefits get described in words (“faster”, “less drudgery”) while the costs arrive on an invoice: a bill per seat or a bill per use, the time spent learning the tool, the mistakes that need correcting. The cost side prices itself. What came back for it does not. Measuring a CRM rollout can fall back on counting, X reports last quarter and 2X now. Nothing here counts that cleanly.
The 3 layers
Section titled “The 3 layers”Output quality. Whether the AI’s work is good enough on its own. The most concrete layer: the model produces a thing, and the thing is either right or it isn’t.
Process change. Whether the work now happens faster or differently. Each individual output can be unremarkable while the workflow around it has genuinely changed.
Outcome impact. Whether the business results are better. Did the AI move a number the business actually cares about?
Doing all 3 is rare. Doing none is common. Doing one of them well is worth more than doing all 3 badly.
Layer 1: output quality
Section titled “Layer 1: output quality”This layer asks whether the model’s work is right, and the way to find out is to read it. Before the rollout this layer was a decision about whether to start, taken on past cases with known right answers. After it, the same reading runs on live work, where nobody knows the right answer in advance. A reading of 50 random outputs from the first month, rated against the criteria that matter for the use case, is enough to see the shape of it. The criteria that recur:
- Complete. It covers what was asked.
- Accurate. The facts and claims hold up.
- Appropriately toned. It matches the brand voice and the context.
- Usable as-is. It could be sent without editing.
- Better than the alternative. It beats a junior team member’s first draft.
Three methods do the reading.
Spot checks. Random outputs, pulled regularly and read against the criteria. Patterns of failure surface this way: always wrong on dates, always too long, always too formal. Those patterns are what improve the prompt or the system around it.
Side-by-side comparisons. The same task done twice, once by the AI and once by a person. A third party rates which is better, ideally without knowing which is which. This is the method to reach for when the question is whether AI should be in this workflow at all.
Rubric scoring. A defined rubric, each output scored 1 to 5 per criterion, tracked over time. This is the method to reach for when prompts or models keep changing, since every change turns one of the three knobs and the question is whether the turn was an improvement.
Reading tends to be heaviest at rollout, while the failure modes are still unmapped, light at steady state, a handful of spot checks a week, and heavy again after any change: new prompt, new model version, new vendor.
The instrument for this layer is the test set built before the rollout. After launch it stops being a one-off test and becomes the thing re-run every time something changes, with the cases that go wrong in live use added to it as they turn up. It does not catch everything. It catches the obvious regressions, and it gives one comparable reading across versions. Without it, prompt and model changes are flying blind. The same set does a second job later, where it is what catches a model version that has quietly drifted.
Layer 2: process change
Section titled “Layer 2: process change”Even when the individual outputs are the same quality as before, AI can change how the work happens, and that is often where the real value sits.
Time per task. How long the task took before the rollout, against how long it takes after. Published figures for language-shaped work sit all over the place, which is the reason to measure the one workflow in front of the team rather than borrow a number. Some tasks lose most of the time. Some lose none, and that is also a finding.
Output per period. Same team, more output produced. Same output, less time spent. Either signals real process change.
Volume handled per person. How much one person clears in a working day, before and after. A rise here at the same headcount and the same hours means less handling time on each item.
Reduced rework. AI first drafts that need fewer revisions than human first drafts are a process improvement before any time-saved calculation is run.
Three things distort this layer when it is read at face value.
Hidden costs. A team that looks 30% faster on the part of the workflow that is measurable may be spending a quarter of that gain on the part that isn’t: prompting carefully, reviewing outputs, correcting mistakes. Honest process measurement covers the full cycle, not just the part the AI touches.
Quality regressions disguised as speed. Faster outputs that are worse outputs are not an improvement. Time without quality is half the picture, and the missing half is usually the one that matters.
Metric gaming. A team that knows its AI usage rate is being tracked will run AI over the easy cases, which were never the slow ones, and leave the hard cases alone. Tracking usage as the main number quietly rewards the behaviour that produces the least benefit.
Layer 3: outcome impact
Section titled “Layer 3: outcome impact”The hardest layer to measure honestly. Did the business actually change as a result of the AI being there?
Outcome metrics, by function:
- Sales: revenue per rep, conversion rate, sales cycle length, leads worked per week
- Support: customer satisfaction, first-response time, resolution rate, escalation rate
- Marketing: campaign output volume, response rate, cost per qualified lead
- Operations: cycle time, error rate, on-time delivery
- Hiring: time to fill, candidate quality, screening throughput
Three things make this layer hard.
Confounding variables. Many things change at once. A new manager, a pricing shift, a competitor stumble, a product launch: any of them moves the same metric the AI is being held to.
Outcome lag. Sales effects land in quarters. Customer satisfaction effects land in months. A 6-week-old rollout against a metric that has not moved is either a failure or simply early, and from inside the moment those look identical.
Noise. Quarter-to-quarter variation in most business metrics is large. A real 5% improvement can sit invisible inside the normal range of variation.
Four things still work against all that.
One or two metrics, not ten. A comprehensive scorecard dilutes the signal. Picking the one or two that matter most for the use case beats tracking a dashboard nobody reads in full.
A baseline recorded before rollout, over at least one or two cycles. No baseline, no comparison. Without a record of where things were before the AI arrived, a claim about what it changed cannot be checked in either direction.
A post-rollout reading with everything else that changed named. If the metric moved, what else moved in the same period? If several things did, the AI’s contribution cannot be cleanly separated out, and saying so is more useful than a confident attribution that won’t hold up.
A split test where the volume allows one. Half the team uses AI, half does not, and outcomes for both groups are tracked over a quarter. This is the strongest method at this layer, because whatever else happened in that quarter happened to both halves. It is also rare, because most companies roll AI out to everyone at once and lose the comparison.
A split test takes discipline: assignment held steady, no spillover between the groups, the same access and training on both sides. What it does not remove is the two problems above it. Sales effects still land in quarters, and against a metric that swings 10% on its own, a real 5% gain needs a lot of leads on each side before it shows through. Where the volume is there, one quarter of a clean split settles what 6 months of dashboard-watching cannot.
What looks measurable but isn’t
Section titled “What looks measurable but isn’t”AI usage. How often the AI is being called says nothing about whether the work is better. A team can use AI heavily and produce nothing of value. Vendor dashboards full of AI activity metrics make leadership feel informed without informing anyone.
Team satisfaction. A team can genuinely enjoy a tool that is not helping them produce more or better work. A team can also resent a tool that is measurably making them faster. Satisfaction is worth knowing and it is not the answer.
Token volume or request count. Measures of cost, not of value. Useful inside an infrastructure conversation, misleading inside an impact conversation.
Number of features used. Having more AI capabilities in the stack does not mean more value delivered. This is a vendor-side number that has wandered into customer dashboards.
Vanity dashboards. Any dashboard where the numbers look impressive and none of them trace to a business outcome is measuring engagement, not impact. They tend to grow where measurement is delegated and nobody is held to outcomes.
What a responsible setup costs
Section titled “What a responsible setup costs”For any AI tool in production, the floor is small. The test set the tool was tested on before launch, re-run regularly with the results kept. A rough before-and-after reading of time per task or output per period, where back-of-a-napkin is acceptable and absent is not. One or two outcome metrics that matter for the use case, with a baseline recorded before rollout. And a named person who reads the worst outputs, because the rare bad ones matter more than the average good ones and somebody has to be watching. (Building trust covers the review patterns at each level of blast radius.)
None of that needs a dashboard. A spreadsheet updated monthly is more than most companies do, and doing the measurement matters more than the sophistication of whatever does it.
When a project isn’t working
Section titled “When a project isn’t working”The hardest measurement decision is the one to stop. AI delivers in some use cases and not others, and the gap between “we haven’t found the angle yet” and “this isn’t the right fit” is where many AI projects quietly stall.
The honest signals that it is not working in a given case:
- After 2 to 3 months, no clear time or output improvement at the team level.
- The team has quietly stopped using it. This often happens without anyone saying so, and usage data is the only reliable read of it.
- Quality issues need so much review that net time is the same or worse than before.
- Outcome metrics have not moved beyond noise.
- The cost is real and recurring, and the benefit, asked about plainly, is a feeling.
At that point the responsible move is to stop or rescope. Money already spent is not an argument for spending more; the only money still on the table is next month’s, and whatever has gone is gone either way.
Projects get carried a long way past the point where they were still teaching anything. Saying “this isn’t working in its current shape” is rarer than it should be, and worth practising before it is needed. A pilot that runs honestly and comes back no has done its job at pilot cost.
Baseline before, not after
Section titled “Baseline before, not after”The single most valuable measurement habit is to write down the baseline before rollout rather than after.
One line recorded before the AI arrives, “130 tickets a day, 18 minutes average handling time, CSAT 4.2”, is worth more than any dashboard built later. Once the rollout happens the past is gone, and there is no honest way to reconstruct what the workflow used to be.
A 5-minute measurement before rollout settles a lot of 6-month debates about whether AI is working. The debate becomes a comparison instead of an argument.