AI in Business: Measuring the Real Return, Beyond the Trend
Finance departments rarely dispute the value of AI; rather, they dispute the figures presented to them to demonstrate that value—and they often have good reasons to do so.
Yet there is no shortage of measurement tools—Microsoft provides three of them. However, they measure usage, and usage is not the same as return on investment. Confusing the two is like presenting a case that won’t stand up to the first serious review.
What Microsoft Tracks for You
The Copilot control system organizes these capabilities into three areas: security and governance, management controls, and measurement and reporting. This last area is divided into three interfaces, each of which serves a different purpose.
- The usage reports from the administration center provide the following metrics: licensed users, active users, active user rate, breakdown by application, active days per person, and date of last activity. Time frames of 7, 28, 90, or 180 days.
- The Copilot dashboard provides an overview of adoption and user experience across the entire organization. It requires no additional licenses —neither a paid Viva Insights subscription nor a Copilot license—to view it.
- Copilot Analytics provides the section intended for senior management, including a business impact report.
These three metrics deliver on what they promise: they track gestures, and they track them accurately. Microsoft counts only intentional gestures as valid—opening a panel doesn’t count, but submitting a request does. The displayed rate is therefore not inflated.
What Microsoft Doesn't Measure—and Says So
The dashboard displays assisted hours, which are not a measure of time saved. The value is obtained by applying multipliers derived from Microsoft’s research to usage volumes, and Microsoft itself describes these values as broad estimates rather than precise calculations. The derived monetary value then applies a default hourly rate of $72 to these hours, based on a U.S. labor market statistic.
Presenting a profit calculated using a default U.S. hourly rate to a Swiss finance department risks having not only that line item questioned, but also the accuracy of the entire report.
There are, therefore, three things that remain beyond the scope of any report, regardless of the subject matter:
- the actual time saved, by job category and by task;
- the quality of what is produced;
- the cause-and-effect relationship between the tool and a business outcome.
The only method that produces a defensible figure
Microsoft makes no secret of the answer: the business impact report does not predict any results. It requires you to upload your own results data, and it provides a list of the metrics expected for each function.
- Sales: average deal size, customer retention rate, cost per lead.
- IT: Number of resolved requests, resolution time.
- Human Resources: Employee Engagement, Retention Rates.
This requirement is real, and it is the only path that will yield a defensible result. The reasoning to follow is based on experimentation rather than on metrics:
- Choose a role and a metric that it has been tracking for at least one year.
- Determine its value before any deployment. Without a starting point, there is no way to measure it.
- Enable this feature—and only this one.
- Wait two quarters. One quarter isn't enough to filter out seasonal noise.
- Compare, and actively look for other possible explanations before drawing a conclusion.
Our Reading
Our position runs counter to the prevailing view: in most organizations, measuring the return on investment for AI isn't worth the cost.
Implementing the system described above requires an existing metric, a clear starting point, two quarters of patience, and the discipline to refrain from equipping everyone in the meantime. Many companies have neither the metric nor the patience, and so they come up with a cobbled-together figure that they defend poorly, at the expense of their finance department’s trust on a topic that, at its core, made perfect sense.
In that case, we recommend setting aside the proof of value and shifting focus. Focus on cost per active user: it’s a fact, it can be calculated in half a day, and it allows for real decision-making—at this price, for this use case, we either continue or scale back. The goal is more modest than that of a return on investment, but the reliability is incomparable.
Comprehensive measurement is worth the effort in one specific case: when a single function is responsible for the rollout and is already tracking a reliable metric. A customer service department that has been measuring its resolution time for three years can produce credible evidence. A company that has distributed licenses to everyone will not be able to do so, because the effect is diluted to the point of being undetectable.
For those who have already rolled out the system on a large scale, it’s not too late—as long as you’re willing to narrow the scope of measurement. Pick one function, isolate it, and measure that one. You won’t get the overall figure you’re asked for, but you’ll get an accurate number—and that’s the one that drives a decision.
What to Do
- Decide now whether you’re aiming for proof of value or cost per use. Both approaches are valid. Mixing the two is not.
- If you're looking for proof: choose the feature and metric before issuing a single license, and note the starting point.
- Never present the hours of assistance shown on the dashboard as a measured result. Call them what they are.
- Reject the default hourly rate. If you need to use it, replace it with your own.
- Present the 90-day active user rate at each committee meeting. That’s the number that sparks the right questions.
- Wait two quarters before drawing any conclusions.
Setting up this system requires no tools other than those you already have; it requires choosing the metrics and sticking to the process for six months. Lambert Consulting can set up this measurement framework with the relevant features and honestly determine whether it’s worth implementing in your organization. Practices that yield results as early as the first month and the identification of candidate processes are the two most common entry points to this question.
Microsoft Sources
What an article Can't Know
An article describes what applies to everyone. What varies from one organization to another is the inventory: which applications, which accounts, and which pieces of equipment are actually involved in your organization. The inventory determines the scope of the effort, and it cannot be summarized on a single page.
You'll be speaking directly with the engineers who will be doing the work, not with a middleman. We'll respond within 24 business hours.
Check what is still true
Announced dates are sometimes postponed, products are renamed, and conditions change. The blog tracks these topics over time: when a rule changes, a new post announces it.
Search for a topic in the blogIn the same issue
Three articles on the same topic. The blog has 149 articles, all of which are freely accessible.

