I’ve been sitting with this one for a while and still don’t have a clean answer.
On a team where AI use is uneven, the people leaning harder on it produce more visible output. That’s measurable. They ship more code, more memos, more analysis, more slides. Their quarterly dashboard looks good. Their velocity looks better than it did a year ago.
The people using it less produce less visible output but may be developing more underneath it. If what I’m seeing and reading is right, they retain more of the learning that used to sit inside the first pass of the work. Slower in output terms, possibly faster on the underlying skill curve. The heavy users are faster now, but some of the learning may be flattening because the model is doing the part of the work where that learning used to happen.
If I’m the manager, I’m running two problems at once. In any given quarter, the heavy users look like my top performers. That’s what the visible read says. If I promote on that read, I may be selecting for output in this stretch of the transition and against the ability to do the work when the tool fails or gets cheap enough that everybody has it. If I weight the slower users for their learning, I’m betting on a signal I can’t measure cleanly against people whose delivered output is lower and whose managers one level up may not see the bet the same way I do.
Things I’ve ruled out. ‘Measure both’ doesn’t work, because the measures pull in opposite directions and whoever sits in the calibration meeting ends up choosing one. ‘Let the market decide’ is a dodge that pushes the problem onto the people themselves. ‘Score on proficiency with the tool’ rewards enthusiasm more than it rewards skill, and the early adopters get counted twice. ‘Wait and see’ accepts drift in the meantime, and two years of promotion decisions come out of that drift.
What I’ve tried is separate tracks for six months, heavy users assessed on output, lighter users assessed on judgment. It creates its own problems. The team notices, the tracks become labels, the labels become political, and the original measurement question is buried under the politics by month four. I still don’t know whether the problem with that approach is the approach itself or just the fact that everyone can feel how provisional it is.
If you’ve solved this at real scale, I want to know what held up once promotion and calibration got involved.