"Reviews measure performance."
One awkward conversation a year, pretending to capture twelve months of work.
Lie No. 09 — Reviews measure performance.
The Myth
he annual performance review is how we objectively measure contribution. Sit everyone down once or twice a year, score them against a scale, rank them against each other, and you get a fair, accurate reading of who performed and who didn't — the data on which pay, promotion, and dismissal can be fairly decided.
The Cold Truth
It's a guessing game dressed as measurement. Once-a-year feedback collapses into recency bias — one bad week sinks a great year (judging a marathon by the final 100 metres). The bell curve / forced ranking manufactures losers and is statistically wrong: real performance is a power law (Aguinis), not a tidy normal distribution. And it's subjective — favouritism and affinity bias make "objective evaluation" a mirage.
The fix everyone reached for was data — and the AI era took it to its conclusion (software scoring performance continuously, AI writing the reviews, firms grading "AI-driven impact"). It scaled the problem: Goodhart's Law means metrics get gamed, AI optimises what's easy to measure not what matters, and a review written by a machine is the opposite of a human paying attention.
The Reframe
Stop chasing objectivity you can't have — accept that judgement is subjective, and design for it. Multiple perspectives, concrete examples, and honesty about uncertainty beat a false-precision number every time.
Separate evaluation from development — the conversation about "here's your rating / here's your pay" and the conversation about "here's how you grow" cannot happen in the same meeting; the moment money is on the table, the development conversation dies. Split them.
Make feedback continuous, not annual — feedback delivered in real time is useful; feedback delivered eleven months late is archaeology. Frequent, low-stakes conversations remove the recency-bias problem and let people actually course-correct.
Kill the forced curve — don't slot a real team into an imaginary distribution. Assess people against the actual requirements of their role, not against a quota that guarantees manufactured losers.
Use AI as an input, not a judge — if AI informs evaluation, keep a human accountable for every decision that affects someone's job, and treat its output as one flawed data point among many, not an objective verdict. Automating bias at speed is still bias.
References
O'Boyle, E. & Aguinis, H. (2012), "The Best and the Rest: Revisiting the Norm of Normality of Individual Performance," Personnel Psychology — individual performance follows a power-law (Paretian), not a normal, distribution.
Goodhart's law — when a measure becomes a target, it ceases to be a good measure. Cognitive bias in evaluation — recency bias, the availability heuristic, and affinity bias as documented distortions in subjective rating.
EU AI Act — AI used for worker evaluation classified as high-risk (Annex III); high-risk applicability date expected to move from August 2026 to 2 December 2027 under the 2026 AI Omnibus (pending formal adoption).