Survey Nonresponse Bias: The Missing Third, and Why Engagement Deltas Are Partly Response-Rate Deltas
A typical enterprise engagement survey closes with a response rate somewhere between 60% and 70%, and the readout proceeds as if the respondents were the workforce. They are not. The missing 30–40% are not a random sample of employees, and when the response rate moves between cycles, a portion of the reported score movement is a change in who answered, not a change in how anyone feels. This piece covers why nonresponse is directional, the minimum defensible correction, and the limits of what any correction can recover.
Nonresponse is not missing at random
The survey-methodology literature distinguishes between data missing completely at random, missing at random conditional on observed variables, and missing not at random. Engagement surveys almost never meet the first condition. Response propensity varies systematically with attributes the organisation can observe:
- Tenure: very new and very long-tenured employees respond at different rates from the middle of the distribution.
- Work pattern: frontline, shift-based, and field populations respond at lower rates than desk-based staff, partly through access and partly through time.
- Manager behaviour: teams whose managers promote the survey respond at higher rates, and those managers are not a random subset of managers.
- Disengagement itself: the employees least invested in the organisation are plausibly the least likely to spend twenty minutes telling it so.
The last mechanism is the uncomfortable one, because it is not observable. It means the respondent pool is likely tilted toward the engaged, and the published favourability score is likely an overestimate of the population value by an unknown amount.
How response rates manufacture deltas
Consider a unit whose response rate rises from 55% to 75% after a participation drive. If the additional 20 points of respondents are drawn disproportionately from lower-engagement groups, the unit’s mean favourability can fall even though no individual’s sentiment changed. The reverse also holds: a cycle with weaker participation can show an “improvement” produced entirely by a more selective respondent pool.
To make the arithmetic concrete — numbers constructed for demonstration, not measured: suppose respondents in cycle one score 72% favourable and the newly recruited respondents in cycle two score 58%. Holding everyone else constant, the unit’s headline falls by roughly four points. A programme owner will read that as a decline. It is a composition shift.
The practical rule is that a score delta is uninterpretable without the response-rate delta and the respondent-composition delta beside it.
The minimum defensible fix: weighting
Two techniques form the baseline, and they are complementary.
Post-stratification reweights respondents so that the respondent pool matches the population on known strata — typically job family, level, tenure band, location, and work pattern. Each respondent in a stratum receives a weight of population_share / respondent_share. The precondition is that the organisation can attach HRIS attributes to responses at a level of detail that confidential reporting permits, which usually means weighting is done centrally before any cut is published.
Response-propensity weighting models the probability of responding as a function of observed attributes, typically with a logistic model fitted on the full population with responded as the outcome, and weights each respondent by the inverse of their predicted propensity. It handles many covariates at once without requiring every stratum to be populated, which matters in small units.
Three operational disciplines go with either method:
- Report weighted and unweighted together. A large gap between them is itself a finding about the respondent pool.
- Trim extreme weights. A handful of respondents carrying weights of 8 or 10 make the estimate unstable; cap and document the cap.
- Report effective sample size. Weighting reduces precision. The effective
nafter weighting, not the raw respondent count, governs whether a cut is publishable.
What weighting cannot save
Weighting corrects for imbalance on observed variables. It cannot correct for selection on the outcome itself. If, within a tenure band and job family, the disengaged are less likely to respond than the engaged, post-stratification on tenure and job family leaves that bias fully in place. No reweighting scheme can recover the opinions of people who systematically decline to give them.
Nor can weighting rescue very low response rates. When a stratum has a handful of respondents, weighting amplifies whatever those few people said. Below a threshold — set in advance, not after seeing results — the right output is suppression.
What we cannot claim
We cannot say how large nonresponse bias is in any given survey, and neither can the survey’s owners. The defensible posture is to bound it: show the response rate by stratum, show weighted against unweighted results, and treat any cycle-on-cycle delta smaller than the plausible composition effect as noise until proven otherwise. A survey that reports its response composition is a measurement instrument. A survey that reports only its scores is reporting the opinions of whoever turned up. The data always wins over the narrative — but the data includes the people who did not answer.