Time saved is only one part of AI ROI. If you only measure hours, you risk declaring success for projects that do not change business results—and abandoning ones that are creating real value.
Why measuring only hours is not enough
Many companies count minutes saved on a task and stop the evaluation there. It is a useful start, but incomplete.
According to McKinsey – State of AI 2025, only 39% of organizations report an EBIT impact at enterprise level, and most of those speak of a contribution below 5%. At the same time, the high performers (about 6%) who see significant value combine efficiency with growth and innovation goals.
The risk is clear: projects that “save time” but do not improve quality, error reduction, satisfaction, or revenue stay invisible—or are judged failures for the wrong reasons.
A five-layer framework (simplified for SMEs)
A practical approach separates five measurement layers:
- System health – availability, latency, technical error rate, cost per query/token.
- Output quality – accuracy, source citability, human correction rate.
- Adoption – active users, usage frequency, share of tasks that go through the tool.
- Process impact – cycle time, errors, bounce-backs, throughput.
- Economic impact – avoided costs, enabled revenue, reduced risk, margin.
You do not need to measure everything from day one. You need to choose 2–4 metrics per layer and link them to the use case.

Concrete metrics for typical use cases
Document / knowledge search
- % of citable and verified answers
- Average time to find the correct information
- Rate of “I have to ask a colleague”
Ticket or email classification / routing
- Classification accuracy
- Average handling time
- % of tickets resolved on first pass
Report or summary generation
- Production time
- Correction rate
- Actual use of the report by decision-makers
Customer care support
- Average response time
- Related CSAT / NPS
- Escalation rate
Anonymized examples: numbers to use as a baseline
These do not replace a full case study. They fix before/after orders of magnitude, aligned with typical Zendata projects and market benchmarks.
1) Knowledge assistant (entity / company with wikis and procedures)
| Indicator | Typical baseline | Credible target at 60–90 days |
|---|---|---|
| Time to find the correct info (recurring questions) | 15–30 min | 2–5 min |
| Answers with a citable source | not systematic | >70% of useful interactions |
| Target users active / week | low or Shadow AI | 30–50% if the tool is in the work channel |
| Recovered time per active user | — | often 40–60 min/day in enterprise AI contexts (OpenAI 2025) |
Reference corpus in real projects: thousands of pages (also 6,000+). Without a baseline on time and citability, “we have a chatbot” is not ROI.
2) Invoice/order reconciliation (feasibility on ~16,500 documents)
| Indicator | Value observed in feasibility |
|---|---|
| Match with rules alone | ~41% |
| Match with rules + AI | ~61% |
| Estimated coverage with assisted proposals + document search | 70–74% |
| Accuracy with a readable order reference | ~96% |
Here ROI translates to: back-office hours avoided × fully loaded cost, minus human review time on exceptions. If 30–40% stays in the human queue, that is not a failure: it is the correct design (human-in-the-loop).
3) Training content from slides/documents (typical pilot)
Typical perimeter: 2 courses of a few hours each, with copy, audio/video, and tests. Sensible metrics: instructional-design hours reduced (often 40–60% on the first draft), SME correction rate, time to publish. Value is not “video generated”: it is a shorter production cycle at domain-accepted quality.
Without a baseline and without ownership of measurement, any percentage is marketing.
Three frequent mistakes in measuring ROI
- Counting only direct costs and forgetting training, integration, monitoring, and opportunity cost.
- Measuring too early (before adoption stabilizes) or too late (when the pilot is already dead).
- Having no clear baseline before introducing the tool.
Without a baseline and without ownership of measurement, any number is debatable.
How to start in practice (also in an SME)
- Choose one process.
- Define 3–5 metrics (at least one for quality, one for adoption, one for process).
- Collect the baseline for 1–2 weeks.
- Launch the pilot with the same metrics.
- Review at 30 and 60 days: if there is no observable movement, revise or stop.
ROI is not a magic formula. It is the discipline of linking the tool to results the company recognizes as important.
FAQ
What is a good horizon to see the first numbers?
On a narrow scope, 30–60 days can be enough for adoption and quality signals. Clear economic impact often needs more cycles.
Do you need a data scientist to measure?
No. Simple metrics, a spreadsheet or light dashboard, and clear process ownership are enough.
What if ROI does not show up?
Check adoption, source quality, tool usability, and whether the chosen metrics are truly linked to value. Sometimes the problem is not AI, but the use-case design.
Does time saved not count?
It counts, but only if it turns into real capacity (more quality work, fewer errors, more attention to customers) and not into time scattered elsewhere.
Sources
- McKinsey: The State of AI (Global Survey 2025): EBIT impact and high-performer profile.
- McKinsey: From promise to impact – five-layer measurement framework: layered framework.
- OpenAI: The State of Enterprise AI 2025: typical recovered time for enterprise AI users.
- Istat: Enterprises and ICT 2025: Italy adoption context.
Dig deeper in the series
- From pilot to production
- Where to start with AI in the company
- Not which AI to choose: which knowledge to make usable
If you want to define metrics and baseline for a concrete use case, we can run a 90-minute session on process, indicators, and measurement plan. Write to info@zendata.it or visit zendata.it.
Pietro Ciattaglia, CEO of Zendata AI, Rome

