In 1996, Avraham Kluger and Angelo DeNisi published a meta-analysis covering 607 studies on feedback interventions — more than 23,000 observations accumulated over four decades of research. The finding that entered the literature and stayed there was this: 38% of the feedback interventions they examined actually reduced the performance of the people who received them. More than one in three. The feedback made things worse.
The researchers were not studying malicious or incompetent delivery. They were studying the ordinary feedback that managers in ordinary organisations give to the people who report to them. The annual performance review, the mid-year check-in, the after-action conversation in the manager’s office. And in more than a third of cases, the person performed worse afterward than they had before the feedback was given.
This is the empirical starting point for any honest conversation about performance management. The intervention organisations invest the most time and emotional energy in — the formal assessment of how a person is performing and what they need to do differently — is demonstrably counterproductive a substantial fraction of the time. The question is why, and the answer has direct consequences for how the executive who leads other executives should think about the feedback they give and the culture of assessment they build.
What Kluger and DeNisi Actually Found
Kluger and DeNisi’s analysis identified the mechanism behind the counterproductive effect. When feedback directs the recipient’s attention to themselves — to their perceived ability, their standing relative to others, their identity as a competent or incompetent performer — it activates a self-protective orientation that is incompatible with the learning the feedback was meant to produce. The person is no longer attending to the task. They are attending to themselves. And attending to themselves produces the cognitive and physiological correlates of social threat: elevated cortisol, narrowed attentional focus, reduced working memory capacity, and the kind of defensive reasoning that produces rebuttal rather than reflection.
The feedback that produced performance improvement in Kluger and DeNisi’s dataset was feedback directed at the task — specific, actionable information about what the work contained and what a better version of it would look like. The feedback that produced the performance decline was feedback directed at the person — evaluations of the individual’s competence, character, or standing in the organisation’s assessment of them. The distinction is not one most performance management systems make. The annual review format, in particular, systematically produces the second type while believing it is producing the first.
What the Physiological Research Adds
The threat response that Kluger and DeNisi’s mechanism implies has a precise physiological substrate. When an individual receives evaluative feedback that threatens their sense of competence or standing — and the annual performance review, by design, positions the recipient as the object of another person’s assessment of their value — the HPA axis activates at levels that directly impair the functions that genuine feedback processing requires. Working memory narrows. The prefrontal circuits that support perspective-taking and non-defensive reception of difficult information are competing for resources with the amygdala-driven self-protective response that the social threat has triggered.
The practical consequence is that the most important feedback — the observations that require the recipient to genuinely update their model of their own performance — arrives in exactly the physiological conditions that make genuine updating least likely. The person listens. They may take notes. They may nod and indicate understanding. The update does not occur at the rate the giver of the feedback believes it has occurred, because the cognitive architecture required to receive and integrate feedback honestly is not fully available in the state the feedback interaction itself has produced.
This is not a criticism of the people involved. It is a description of what human nervous systems do in evaluative social contexts — which is the same thing they do in any context where social standing is being assessed by someone with power over the outcome.
What Effective Feedback Actually Looks Like
John Hattie and Helen Timperley’s 2007 review of the feedback literature distinguished between four levels at which feedback can operate, and found that these levels differ substantially in their effects on learning and performance. Feedback at the self level — the kind that tells people about their identity and worth as a performer — produced the weakest effects and, consistent with Kluger and DeNisi, was often negative. Feedback at the task level — specific information about the correctness, quality, or accuracy of the work — produced moderate positive effects. Feedback at the process level — information about the strategies and approaches the person used, and what alternative approaches might work better — produced stronger positive effects. Feedback at the self-regulation level — observations that help the person develop their own capacity to monitor and adjust their performance — produced the strongest effects of all.
The hierarchy matters because it maps directly onto what most executive feedback conversations actually contain. The typical performance discussion allocates most of its time to the task level at best — observations about specific recent outputs — and slides easily into the self level when the content is critical. The process and self-regulation levels, which are where the research-supported performance effects are strongest, are the levels that require the most psychological safety to access honestly, because they involve the executive making observations about how the person thinks and approaches their work — which is the category of feedback that feels most evaluative and activates the threat response most reliably.
The Annual Review Format Specifically
The annual performance review aggregates feedback across a twelve-month period and delivers it in a single conversation. This format has a specific failure mode that the research supports: the combination of retrospective aggregation — which is highly susceptible to recency bias, halo effects, and the manager’s current relationship quality with the person — with high-stakes delivery, in which the conversation outcome affects compensation, promotion, and role security, produces exactly the evaluative self-threat conditions that Kluger and DeNisi found to be counterproductive. The person enters the room knowing the conversation will contain an assessment of their standing, which activates defensive cognition before the first word of feedback is delivered.
The format also produces what researchers have called the feedback sandwich — the practice of bracketing critical observations between positive ones to soften the threat response. The evidence for this technique’s effectiveness is weak, and there are reasonable grounds to think it is actively counterproductive: recipients become skilled at identifying the structure, which makes the positive observations read as throat-clearing for the criticism rather than genuine recognition, and the critical observation, having been identified, activates the same threat response it would have activated without the sandwich. The sandwich changes the sequence. It does not change the physiological effect.
What a Feedback Practice That Works Requires
The research supports a set of conditions that are collectively unfamiliar in most organisational feedback cultures. Feedback should be proximate to the work — delivered close to the moment of the behaviour rather than aggregated across a year. It should be specific to the task and process rather than evaluative of the person. It should be delivered in conditions where the recipient’s threat response is not already elevated — which means separating developmental feedback from compensation and promotion conversations rather than combining them in the same interaction. And it should be bidirectional: Edgar Schein’s research on diagnostic versus pure inquiry suggests that the most useful feedback conversations are ones in which the manager is genuinely curious about the person’s own assessment of their work, not simply delivering a pre-formed conclusion.
None of this is technically difficult. All of it requires the executive giving the feedback to have sufficient physiological regulation to hold the conversation without either avoiding the difficult observations or delivering them in a way that activates the recipient’s threat response. That regulation is not a communication skill. It is a physiological state — and it is the precondition for the kind of feedback that actually changes performance rather than simply satisfying the organisation’s administrative requirement to document that the conversation happened. Four slots available monthly. Apply here.
Frequently Asked Questions
If feedback is so often counterproductive, should organisations stop giving it?
The Kluger and DeNisi finding is not that feedback doesn’t work. It is that feedback directed at the person’s identity and standing rather than the task and process doesn’t work, and that the typical delivery format triggers the self-protective response that makes the feedback ineffective. The 62% of interventions in their dataset that did not reduce performance — many of which produced genuine improvement — were almost uniformly task-directed, specific, and delivered in conditions where the social threat element was low. The implication is not to stop giving feedback but to change what feedback conversations contain and how they are structured. Organisations that replace the annual review with more frequent, task-specific, bidirectional conversations consistently report both higher engagement and better performance outcomes than those that retain the traditional format — which is consistent with what the research would predict.
What is the relationship between psychological safety and effective feedback?
Amy Edmondson’s psychological safety research and the Kluger and DeNisi feedback findings are describing the same dynamic from different directions. Edmondson’s work establishes that people in teams with high psychological safety are more likely to speak up, admit errors, ask questions, and engage with difficult information — because the social cost of doing so is lower. The feedback research establishes that people who receive feedback in high-threat conditions process it defensively and update their performance less. Both lines of research converge on the same architectural point: the quality of cognitive and behavioural engagement with difficult information is a function of the perceived threat level of the social context. Building the conditions for effective feedback and building the conditions for psychological safety are the same project.
How do high-performing organisations handle this differently in practice?
The most consistent differentiator in organisational research is frequency and separation. Organisations with strong developmental cultures tend to separate the compensation conversation from the performance conversation — holding them at different times with explicitly different purposes — because combining them ensures that the developmental conversation will be perceived as evaluative regardless of how it is framed. They also tend to hold feedback conversations more frequently, which keeps the observations specific and proximate rather than aggregated and retrospective. The pattern across high-performing organisations consistently emphasises frequency, specificity, and the structural separation of development from evaluation.
How does the SEAM protocol approach executive feedback and performance development?
The Clarity Index assessment that begins the SEAM protocol is itself a form of feedback — and it is designed with the Kluger and DeNisi findings explicitly in mind. The assessment is framed around the executive’s own data: their HRV baseline, their self-reported capacity across the Clarity Index dimensions, their performance patterns across different times of day and types of demand. The observations that emerge from it are specific to the executive’s physiological and behavioural patterns rather than evaluative of their worth as a leader. This framing matters because it positions the executive as someone using information about their own system rather than receiving a verdict about their adequacy — which produces a fundamentally different receptive state. The 90-day protocol uses the same structure: progress observations are specific and task-directed, and the executive’s own measurement data provides the primary reference rather than the consultant’s assessment of them.