The metric that feels like progress
Every AI rollout eventually produces a dashboard, and the number on it is almost always some form of usage. Tokens consumed, active seats, messages sent, percentage of engineers with the assistant enabled, count of AI-authored pull requests. The number goes up and to the right, which feels like progress, and it gets reported upward as if it were the return on the investment. It is not. Usage is the easiest thing to count and the least informative thing to know.
Usage answers the wrong question
Usage tells you that a tool was used. It does not tell you that the tool helped. Those are different claims, and the gap between them is where AI budgets go to die. A high token count is consistent with a team getting dramatically more done. It is equally consistent with a team generating more output that then needs more review, more rework, and more cleanup than it saved. The counter cannot tell the two apart, because it is measuring activity, and activity was never the goal.
This is not a new mistake. It is the same one teams made counting lines of code and commits, dressed in newer clothes. The measure is seductive precisely because it always moves. Give people a tool and usage goes up. That reliability is what makes it feel like a signal and what makes it useless as one.
Measure the outcome, not the activity
The honest question is whether the work that matters changed. That pushes you back onto the outcome measures you already trust, the ones that were true before AI and stayed true. On the delivery side, that includes whether you are shipping faster without shipping more defects. Change failure rate is the one to watch, because the failure mode of AI-accelerated work is producing more, faster, with more mistakes riding along. If throughput rose and change failure rate rose with it, you did not get faster. You got a bigger pile.
Then ask the question usage can never answer: what work stopped happening, and was it the right work to stop. The point of the tool is not to add activity on top of everything people already did. It is to remove or compress work that used to cost time, and free that capacity for something higher value. If nobody can name what stopped, or the freed time went straight into producing more of the same low-value output, the tool changed the inputs and not the result.
The trap on the other side
There are two ways to get measurement wrong, and vanity metrics are only one of them. The other is to measure nothing, to wave off the question because knowledge work is hard to quantify and run the whole program on vibes and anecdote. That fails differently and just as badly, because you cannot tell a real gain from an expensive habit, and you keep funding both.
The over-correction from there is to build an elaborate ROI apparatus, instrumenting every task and attaching a dollar figure to every saved minute, which costs more to run than it ever returns in insight and produces numbers precise enough to argue about and too synthetic to trust. The useful position sits between them. Track a small number of outcome signals you already believe, connect the usage to whether those moved, and resist the urge to manufacture precision you do not have.
A short way to tell if AI created value
Three questions get you most of the way, without a new measurement program:
- Did a delivery or business outcome you already track actually move, and did it move without a matching rise in failures or rework.
- Can you name the work that stopped happening, and did the freed capacity go to something higher value rather than just more output.
- If you turned the tool off tomorrow, would you feel the loss in a result, or only in the usage chart.
If the only honest answer is that the usage chart would drop, you have measured a habit, not a value.
Where usage metrics are fine
None of this makes usage useless. During an initial rollout, reach is a legitimate thing to track, because before a tool can create value it has to actually be adopted, and usage tells you whether that is happening. The error is not looking at usage. It is calling it value, and letting the adoption number stand in for an impact the tool has not yet been shown to produce. Track usage as adoption, label it as adoption, and do not let it graduate into an ROI claim it cannot support.
Where people will push back
"Usage is a leading indicator of value."
It can be, but only if you eventually connect it to a lagging outcome. A leading indicator you never tie to a result is not leading anywhere. It is a number that reliably goes up. The discipline is to treat usage as a hypothesis about value and then check whether the value showed up. If you never run the second half, you do not have a leading indicator. You have a comfort blanket.
"You cannot measure knowledge-work productivity."
Not perfectly, and you do not need to. The claim that it is unmeasurable usually smuggles in the conclusion that you should stop trying, which leaves you measuring activity by default, which is worse than an imperfect outcome measure. You can detect whether the outcomes you care about moved. That is enough to steer by, and it is more than a token count will ever give you.
"Leadership wants the adoption number."
Then give it to them, clearly labeled as adoption. The problem is never that the number exists. It is when the organization starts celebrating the token count as if it were impact, and makes decisions as though usage and value were the same thing. Report the reach and the outcome side by side, so the activity is never mistaken for the result.
The takeaway
AI usage is easy to measure, which is exactly why it fills dashboards and why it misleads. It proves the tool was used, not that it helped, and those are not the same claim. Usage is not value. It is the cost you are hoping turns into value. To know if it did, look at whether the outcomes you already trust actually moved, and at what work stopped and whether it should have. Everything else is a number that goes up on its own.
Further reading
- AI Isn't Replacing Engineers. It's Rewriting the Org Chart. Why change failure rate is the metric that matters when execution gets cheap.