Did They Get Better, or Were They Always Good?

We measure our AI program carefully, and the number we report is hours saved. Then I looked at what that number is made of, and at who is actually using these tools.

AI adoptionmeasurementoperator notesOneDigitalAI workforce
Two sheets of translucent tracing paper on a warm off-white surface, each bearing a simple hand-drawn rising line, laid one over the other so the two curves read as a single line.

A leader at my company asked me to write down how to do my job without me, for clients who will never have me. I wrote fifteen steps in an afternoon. Ten of them transfer cleanly to anyone willing to do the work, and most of what they describe is how we hire, manage, and occasionally fire these things, which I have written about already.

The interesting part was the step I could not write.

Step five. Before you build anything, decide how you will know the business actually got better. Not that the work got faster. That the company is better off.

I have been reporting an input and calling it a return

I want to be careful here, because the easy version of this admission is false modesty, and it would also be untrue.

We do have a methodology. We built a measurement system for the program. It runs across every conversation our people have with our AI coworkers, it produces hours saved and dollars saved, and I would defend its arithmetic in a room full of people who wanted to poke at it. I have written before about how hard it is to measure good AI use, and that system is my team's best answer to it.

Now look at what the number is actually made of. It takes an estimate of the minutes a given type of task normally consumes, and multiplies it by a burdened hourly rate.

Sit with that for a second, because it took me longer than it should have. The same task, performed by a more expensive person, books more value. Change nobody's behavior at all, shift the same work to a higher-paid part of the firm, and our value number goes up. It is a faithful measure of what the time we saved would have cost us if we had spent it.

It is not a measure of whether we won more business. Or served a client better. Or kept an account we were on track to lose. Or made a consultant's judgment sharper.

Hours saved is an input. I have been reporting an input and calling it a return, and doing it sincerely, which is the part that bothers me.

That would be a tolerable accounting problem if usage were free. It is not. Every consultant working with a coworker spends tokens, and that line item grows at exactly the rate of the adoption number I keep presenting upward as a success. There's a word doing the rounds for optimising the number you can see. Tokenmaxxing. Not mine, and I don't love it, but I've been close enough to it that I've stopped laughing.

The two questions, and the one that keeps me up

The first is attribution. Our people use these tools heavily. Are they winning more business because of it?

The second is harder. The consultants using our coworkers most heavily. Were they already our strongest performers? Or are they people who were behind, and got better?

Those two possibilities produce an identical adoption chart and opposite conclusions.

If the heavy users were already the best, we've only bought an expensive amplifier for the people who needed one least, while the people who needed it most are absent from the data entirely.

If instead the heavy users are people who were struggling and are now closer to the front, that is a different kind of company than the one we were two years ago.

What the research says, and why it raises the stakes

The best available evidence points hard at the second story. In a lab, though.

Researchers at Harvard ran 758 Boston Consulting Group consultants through realistic consulting tasks, randomly assigning who got access to AI tools/agents. The consultants using it did more work, faster, at higher quality. But the finding that matters here is the distribution. Consultants who had scored below the average improved by around 43 percent. Those who had scored above it improved by around 17. The tool compressed the skill gap. It took what the strongest people already knew and handed it to everyone else.

The same shape shows up elsewhere. Another study focused on customer support and found a 14 percent gain overall, concentrated among the least experienced agents.

Two studies, different populations but same direction: AI lifts the bottom more than the top.

But notice the word doing the work in both. Assigned. In each study, researchers decided who got the tool, which is exactly what makes the results trustworthy. Nobody assigned my colleagues anything. They opted in.

A row of small identical wooden pegboards on a warm off-white wall, most holding a single neatly hung coral-handled tool, three of them completely bare.

So I went and looked at ours, and I was wrong

I had a tidy theory about what opting in selects for. Curiosity, confidence with new software, enough slack in your week to look clumsy for an afternoon. All of which, I assumed, correlate with already being good at your job.

Our own data does not support that.

The strongest predictor of whether someone here uses an AI coworker turns out to have almost nothing to do with the person. It is whether anyone has an AI coworker tailor-built for their job. Parts of the business with a dedicated one adopt at roughly double the rate of parts without. Every genuine dead zone in our numbers is an area where we simply never built anything.

The second strongest predictor is their manager. If your manager is a sustained user, you are far more likely to be one, which is the change-management argument I made a while ago showing up in our own telemetry, and that holds inside every practice we tested, which rules out the comfortable explanation that our strong practices happen to have strong managers.

Another finding in my data that vexed me - our newest hires adopt markedly less than our longest-serving people. When my gut told me that it would be easier to get new joinees to adapt to our way of working.

None of that is about individual talent. It's about supply and social proof.

Which means the most interesting group in our data is not who I thought. The people who have never logged in are not our laggards, our sceptics, or our dead weight. Overwhelmingly, they are people nobody built anything for. They are not un-curious. They are un-served. That is a far more uncomfortable finding, because it is our fault rather than theirs, and it is also far more fixable.

These are important findings I can action on. However, it still does not answer my question. A small minority of our users accounts for most of our measured value, which sounds like a finding until you remember what measured value is made of: volume, task type, and hourly rate. It is not evidence that our heavy users improved, and it is not evidence that they were strong to begin with. The instrument cannot see the difference, and that is the whole problem.

And heavy use is not automatically good either

There is a second finding in that same Harvard study I've been thinking about. Outside the range of tasks the AI handled well, consultants using it performed about 19 percentage points worse than consultants working alone. And the boundary is jagged, so you can't tell from a task's apparent difficulty which side of it you're on. So a consultant using a coworker constantly might be compounding a real advantage, or might be producing worse work faster with more confidence. Volume cannot distinguish those two. Neither can a satisfaction score, which is the metric most programs lean on hardest because it is the easiest one to collect.

This is why the most bureaucratic-looking thing we do turns out to matter most. When we write a job description for a coworker, we write down what it will not do. That looks like paperwork. It is us drawing that jagged boundary by hand, one job at a time, because nobody can hand us the map. I didn't plan on it, but that rigor has definitely helped ensure we're going down the right path.

Why the studies will not rescue you

You might reasonably start by looking up what others are doing.

The most quoted statistic in enterprise AI is that ninety-five percent of pilots deliver no return. You have seen it. It usually arrives rendered as "95 percent of AI fails."

Go and read the actual report, because it is considerably more careful than its reputation. It reviews more than three hundred publicly disclosed AI initiatives, runs structured interviews with representatives from fifty-two organisations, and surveys a hundred and fifty-three senior leaders. And it is precise about what it counts. The five percent applies to custom and task-specific enterprise tools, where the funnel runs from half of organisations investigating, to a fifth piloting, to one in twenty reaching production. General-purpose tools like ChatGPT and Copilot run an entirely different course in the same report: over eighty percent explored or piloted them, and nearly forty percent report deployment. No one reports that second number.

The report is equally explicit about what it means by failure. For task-specific tools it defines success as something users or executives remarked had caused "a marked and sustained productivity and/or P&L impact." That is a high bar, honestly stated, and it means a modest genuine return files as a flop.

Then, in the research limitations, the authors write this sentence which makes me chuckle - They say their figures are "directionally accurate based on individual interviews rather than official company reporting," sample sizes vary by category, and, most importantly "success definitions may differ across organizations."

Read that last clause again. The most cited number in enterprise AI, the one everyone reaches for to argue this is not working, carries a footnote from its own authors saying the organisations inside it do not agree on what success means.

That is not a flaw in the research. It is the same wall I hit at step five, turning up in somebody else's data. And it is why the number got flattened on its way to your board deck: every qualifier that makes it true is a qualifier that makes it unquotable.

Why I wanted people who could tell me no

Researchers at Harvard and MIT are working with our data now, and I want to be precise about why, because "we are working with Harvard" is the sort of thing people say to sound impressive (I would be remiss if I said I wasn't proud of that association though!).

The reason is less flattering. A selection problem cannot be settled by anyone with a stake in the answer, and nobody has a bigger stake than I do. I lead the team that built this program. I have spent two years arguing for it inside the company. If I design the study that grades it, I will find what I am hoping to find, and I will do it sincerely, which is exactly what makes it untrustworthy. No amount of personal integrity substitutes for not being the one holding the pen.

The other reason is that this is a genuine open problem rather than a service engagement. Notice what both of those studies had that we do not: a denominator. Resolved tickets per hour. Scored deliverables on an assigned task. A benefits consultant's month does not come with one, and neither does a wealth advisor's or a lawyer's. That is not a failure of rigor. Rigor needs something to divide by, and knowledge work does not supply one.

What the measurement is actually for

I have been describing this as a measurement problem, and it is one. But I do not care about it because I want a better dashboard.

Here is what I actually believe, and I am going to say it plainly, because softening it would be its own kind of dishonesty and because the people it applies to deserve to hear it straight.

AI is going to take a lot of jobs. Not evenly and not randomly. It is going to take them from people who are mediocre at what they do, and from people who refuse to reinvent themselves. Anyone genuinely resistant to change is going to be displaced, and I do not think that is a controversial prediction so much as an unpopular one. I believed it before I ran a program like this. Running one has not softened it.

I am aware of how that sounds coming from the person who runs the AI program at a firm of six thousand people. Which is exactly why the second half of the thought matters more than the first.

If that is true, a company has a choice about what to do with it, and the choice is not really a technology choice. You can let it arrive on its own schedule, notice a year later which roles have quietly become redundant, and handle it the way companies have always handled that. That path requires no conviction and no effort, and I think a great deal of the market is going to take it.

Or you can try the harder thing. Find the people whose work is being automated out from under them before it happens rather than after. Give them the tool, real training rather than a webinar, and a route into work that is worth more than what they were doing. Come out the other side genuinely more productive, with the same people still in the building.

That is the version I want to be able to point at when this is over. Not that we are closing in on seventy percent of our full-time workforce using AI coworkers. That we armed everyone, that nobody had to go, and that we are measurably better off than we were.

None of this is mine alone to worry about. My bosses, Vinay and Mike, have been making the same argument in public and in a book they have coming out: that the real win is keeping your people and turning them into 10x versions of themselves, rather than thinning the payroll in service of the technology. They have put their names to that position, which means they are at least as invested as I am in finding out whether it actually holds.

And here is the uncomfortable part. Everyone in this industry says some version of that sentence. It is easy to say, it costs nothing from a stage, and I have watched people say it who have clearly never checked. The only thing separating a company that means it from a company that says it is whether anyone bothered to find out.

That is what the measurement is for. If we cannot distinguish people who got better from people who were always good, then we also cannot tell whether arming someone changed their trajectory or whether we handed a capable person a faster horse. And if we cannot tell that, then "we reskilled our people instead of replacing them" is just a story we tell rather than a thing we know.

I would rather know. Including if the answer is that it did not work, which is the real reason I want the people doing the checking to be free to come back and tell me no.

Because if the humane version of this turns out to be real, it is worth an enormous amount to prove, to us and to everyone we would then be able to show. And if it is not real, the people most affected deserve to find that out from someone who went looking, rather than from a layoff.

You cannot begin any of that if the only thing you can see is who tokenmaxxed.

Robin's Notebook

A new entry every couple of weeks. No promotion, no funnel, no manifesto. Just the unfiltered version.

Subscribing opens beehiiv.com in a new tab to confirm. Unsubscribe anytime; I won't share your email.