A manager sits down to write a performance review and finds the old template doesn’t cover the work they’re looking at. They can’t tell, from the work alone, how much of the thinking was the person’s and how much came from a model. The output is good, but the review was built for a world where ‘the person did the work’ was a safe assumption, and that assumption isn’t safe anymore.
The review system in most companies was designed around a simple assumption, that the work a person handed in was the work a person had done. Good output meant the person had done the thinking and the writing themselves, and the review could read the output as a reasonable proxy for the effort behind it. The proxy has broken, because a large part of the first draft now comes from somewhere else, and the output no longer carries the same evidence of the work that produced it.
Two ways to handle the shift are common, and both miss. The first is to pretend the AI wasn’t there and review the output on its own, which rewards the person who used the AI most because their output volume goes up for the same hours, and punishes the person who didn’t, whose output looks thin next to theirs. The second is to police AI use and treat the tool as the problem, which produces a company where people use the tool and don’t say so, and a review that is blind to most of the work being done.
The honest version of the review looks at a different layer of the work. The output is part of the evidence, but the larger part is the judgment that shaped it: what the person asked the model, and what they caught in the model’s output that they chose to fix or throw away. That judgment is what the review is supposed to be measuring in the first place, and it’s the thing the AI didn’t do.
In practice the review asks different questions. The old review asked whether the document was good, and the new review asks why this version of it is the one that shipped: what the first draft said, and what the person changed. The artifact becomes the evidence of the decision, rather than the only thing being reviewed. A manager who asks this gets a different reading of the person than the old review produced, because the output was always a lossy summary of the work, and AI has made the loss bigger.
This changes what good work looks like. A person who asked a model for mediocre help and shipped it unchanged did worse work than a person who asked for the same help, read it, noticed three things the model had got wrong, and fixed them before shipping, even though both people spent the same number of hours and both handed in an adequate document. The difference is invisible in the output but visible if you ask about the process.
Most review systems haven’t been updated for this yet, because the update is expensive. The manager has to read the work and know enough about how the person worked to write an honest review of the judgment calls, which is more effort than reading the document and writing a rating. The companies that make the investment get reviews that mean something in a workplace where AI is doing part of the work. The ones that don’t are running a review on a decreasing slice of the work and wondering why the ratings keep drifting away from the real performance.