The value of an AI system is the difference it makes to a specific piece of work, measured against how that work performed before, net of what the system costs to run. That sounds obvious, yet many projects launch without a baseline, report only model accuracy, and leave out the cost of human review. The result is a number nobody fully trusts.
This guide sets out a measurement approach that holds up to scrutiny from finance, operations and the people doing the work.
The short answer
Track four groups of measures, and decide them before you build:
- Outcome metrics that the business already cares about, such as handling time, backlog, turnaround or conversion.
- Quality guardrails that must not get worse, such as error rates, rework, complaints or escalations.
- Unit cost of completing one piece of work, including model usage, infrastructure, human review and maintenance.
- Adoption by the people who are supposed to use the system.
Then compare against a baseline, and be honest about what else changed in the same period.
Start with a baseline
A baseline is a measurement of the current process taken before the AI system is introduced. Without one, any improvement is an estimate.
- Measure the same unit of work. If the system handles invoices, measure cost and time per invoice, not per department.
- Capture variation. Record the spread, not only the average. A process that is fast on easy cases and slow on hard ones behaves differently from one that is uniformly slow.
- Use a representative period. Avoid measuring during a seasonal peak or an unusually quiet month unless that is the period you care about.
- Include the hidden work. Rework, follow-up questions and escalations are part of the cost of the process.
If existing systems do not record what you need, a short period of manual time sampling is better than nothing, as long as you use the same method afterward.
Choose outcome metrics the business already uses
The best outcome metrics are ones that someone in the business already reports on. They carry credibility that a new AI-specific score does not.
| Type of work | Example outcome metrics |
|---|---|
| Document processing | Time from receipt to posting, share of documents needing manual correction |
| Customer service | Time to resolution, contacts resolved without handoff, repeat contacts |
| Knowledge access | Time to find an answer, questions escalated to specialists |
| Forecasting | Forecast error at the level decisions are made, stockouts, excess inventory |
| Software engineering | Cycle time for changes, defects found after release |
Pick two or three. More than that dilutes attention and makes it easy to find a number that looks good.
Set quality guardrails
A system that makes work faster but less accurate has not created value; it has moved cost somewhere else. Guardrails are metrics that must stay at or above their baseline level for the result to count.
- Error rate on the outputs that matter, measured by reviewers who know the work.
- Rework and reversals, such as corrected records or reopened tickets.
- Customer or user complaints related to the process.
- Escalations to specialists or supervisors.
- Policy or compliance exceptions found in review.
Guardrails should be reported alongside outcome metrics, every time. In our view, a value report that omits them should not be accepted.
Calculate the full unit cost
Model usage fees are usually a small part of the cost of an AI system. A fair unit cost includes:
- Model and infrastructure costs, including retrieval, storage and hosting.
- Human review time for outputs that are checked or corrected.
- Exception handling for cases the system cannot process.
- Maintenance, including evaluation runs, prompt and model updates, and integration fixes.
- Support and monitoring effort from the team that owns the system.
Divide by the volume of work completed to get cost per unit, and compare it with the baseline cost per unit. Revisit the calculation as volume changes, because fixed costs spread differently at different scales.
Measure adoption honestly
A system nobody uses has no value, however good its outputs. Adoption measures tell you whether the benefit is reaching the work.
- Active use by the intended group, over time rather than in launch week.
- Acceptance rate of suggestions, and how often people edit them before accepting.
- Bypass rate, meaning how often people route around the system and do the work the old way.
- User feedback, collected in the tool and in short conversations with the team.
Low acceptance is useful information. It often points to a quality problem in a specific category of case, which the evaluation set can then target.
Understand the limits of attribution
This is the part most value reports skip. Other things change while an AI system is being introduced: staffing, volumes, process changes, new policies, seasonal patterns. Not every change in the numbers is caused by the AI.
Ways to strengthen attribution:
- Staged rollout. Introduce the system to one team, region or case type first and compare it with a group that has not yet received it.
- Hold-out cases. Route a small share of work through the old process for a period and compare.
- Before-and-after with context. If a comparison group is not possible, record other changes in the same period and state them next to the result.
- Ranges, not single figures. Present a plausible range of impact, with the assumptions written down.
It is better to report a modest, well-supported result than a large number that falls apart when someone asks how it was calculated.
A simple reporting rhythm
- Before launch: agree metrics, guardrails, cost model and baseline with the business owner and finance.
- First weeks: report weekly on quality and adoption; outcome metrics are often noisy early on.
- After stabilization: report monthly on outcomes, guardrails and unit cost together.
- Quarterly: revisit whether the metrics still reflect what the business needs.
Where to go next
If you need help setting up a baseline and the reporting behind it, see Analytics & BI. To shape the business case before building, see AI Consulting & Strategy.