Select Page

Someone once suggested we reward people based on how many AI prompts they ran. It was well intentioned. It was also one of the fastest “no” decisions I’ve made in this whole program.

I understood the appeal. Adoption metrics were patchy, leaders wanted momentum, and prompt counts are beautifully easy to measure. Set a target, publish a leaderboard, watch the number climb.

And the number would have climbed. That’s exactly the problem.

Reward prompts and you get prompts. People would run queries to be seen running queries. The metric would look magnificent while telling us precisely nothing about whether work was getting better, faster or safer. Worse, it would teach the organisation that AI is about performing activity rather than producing outcomes, a lesson that’s very hard to unteach.

There’s a deeper issue, and it applies well beyond this one suggestion. In the early phase of AI adoption, almost everything easy to measure is activity: licences allocated, logins, prompts, sessions. Almost everything that matters is harder: decisions improved, hours actually returned, errors caught, risk exposures surfaced and fixed. When leaders are anxious for evidence, the easy numbers exert a gravitational pull. Entire programs end up optimised for the dashboard.

What we chose to celebrate instead was real stories that delivered value. A person, a real piece of work, what changed, what it means. The analysis that took half a day now takes twenty minutes, and here’s how I check it before I trust it. That last clause matters as much as the speed. We wanted judgement to be part of what gets applauded, because in our environment an unchecked output isn’t productivity, it’s a liability.

Usage data still has a place. We watch it as a signal, not a scoreboard.

Before you set any AI metric, ask the question that saved us here. If people optimised for this number and nothing else, would we be pleased with what we got?