Define value before measuring automation
Name the outcome in operational or financial terms: cases resolved correctly, time to an approved deliverable, avoided loss, additional contribution or a learning result. State who benefits and over what period. A count of generated messages does not show that the messages helped.
For hard-to-price outcomes, report a balanced scorecard rather than inventing a monetary value. Quality, cycle time, accessibility and risk can be decision-relevant without being forced into one number.
Use a full-cost ROI formula
Total cost can include usage, tools, infrastructure, integration, evaluation, monitoring, review, corrections, training and failure. Separate one-time implementation cost from recurring operating cost so the payback story remains visible.
Match evidence to the claim
| Claim | Evidence approach | Caution |
|---|---|---|
| Work became faster | Comparable cycle-time baseline | Check whether quality or backlog changed |
| Cost fell | Cost per accepted task before and after | Include review and retries |
| Outcome improved | Controlled test or credible comparison | Correlation alone may reflect other changes |
| Risk declined | Defined incident and severe-failure measures | Rare events need longer observation |
Randomised holdouts can support causal claims when feasible. When they are not, disclose the comparison's limits and avoid attributing every change to AI.
Evidence & context: Google Research
Build a small decision dashboard
- Volume attempted and accepted.
- Quality and severe-failure rate by task category.
- Cost and time per accepted task.
- Retry, escalation and human-review rates.
- Business or learner outcome tied to the original objective.
- A baseline, target, owner and review date.
Use the dashboard to stop as well as scale. If value does not exceed total cost at the required quality threshold, redesign the task or retire the workflow.
Sources & further reading
- Methods for Measuring Brand Lift of Online Ads
Google Research. Original research using randomised experiments to estimate advertising effects; no universal lift or ROI benchmark is inferred.
- Generative Artificial Intelligence Profile (NIST AI 600-1)
NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.
- Evaluate a model router
Microsoft Learn. Official guidance for evaluating routing across representative workloads using quality, cost, latency and policy criteria. It is not evidence that routing always improves results.
Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.
A correction, a counterexample or an experience worth sharing?
Join the conversation ↗