Workforce Data Lab
people-analytics Workforce Data Lab · research desk

How Vendor Benchmark Reports Get Constructed — and a Six-Question Filter for Using One as a Planning Input

People analytics teams are asked to benchmark constantly: attrition against industry, time-to-fill against peers, engagement against a norm. The benchmark usually comes from a vendor report, and the number is treated as an external fact. It is better understood as the output of a data-collection process with specific, usually undisclosed, design choices. Two benchmarks of the same metric from two reputable sources can disagree by a factor of two, and both can be honestly produced. This piece describes how benchmark reports get constructed, why they diverge, and the filter we apply before letting one into a plan.

How the sample gets built

Most commercial benchmarks are not samples in the statistical sense. They are aggregations of whatever data the vendor has access to, and the access mechanism shapes the result.

Self-selected submissions. Survey-based benchmarks rely on organisations choosing to participate. Participation correlates with having a mature analytics function, having the data in a reportable state, and in some cases having results the organisation is comfortable sharing. The participating population is not the industry.

Client-mix skew. Platform-derived benchmarks are computed from the vendor’s customer base. That base reflects the vendor’s sales history: certain sizes, regions, and sectors. A benchmark from a platform popular with mid-size technology firms describes mid-size technology firms, whatever the report’s title says.

Undefined denominators. Attrition can be computed against average headcount, opening headcount, or closing headcount; voluntary-only or all exits; with or without interns and fixed-term staff. A benchmark that publishes a rate without its denominator and exclusion rules cannot be compared to an internal metric with a known definition.

Season-of-collection effects. Data collected during a period of high labour-market churn produces higher attrition and time-to-fill norms; a report collected in one window and published a year later carries that window’s conditions forward. Annual reports rarely flag that the underlying period differs from the publication date.

Why two benchmarks disagree by 2×

The divergence is usually definitional and compositional rather than an error in either source. To illustrate — numbers constructed for demonstration, not measured: one report computes total attrition against average headcount across a client base weighted toward high-turnover sectors and arrives at 24%; another computes voluntary attrition against opening headcount among survey participants in professional services and arrives at 11%. Both are labelled “annual attrition.” Neither is wrong. They are not measuring the same thing about the same population.

Divergence of this size should be the expected case, not a surprise. It also means that an organisation can find a benchmark to support almost any narrative by choosing among sources, which is a governance problem as much as a methodological one.

The six-question filter

Before a benchmark is used as a planning input — a target, a budget assumption, an argument for investment — we ask six questions. If more than one cannot be answered from the report, we treat the figure as anecdote.

  1. Who is in the sample, and how did they get there? Self-selection, client base, or a designed sample, with counts.
  2. What is the exact metric definition? Numerator, denominator, inclusions, exclusions, and the period measured.
  3. How does the sample’s composition compare to ours? Sector, size, geography, workforce mix. A frontline-heavy organisation compared to a knowledge-worker benchmark will always appear to underperform.
  4. When was the data collected? Not published — collected.
  5. What is the dispersion? A median without quartiles conceals whether “typical” is a tight band or a wide spread. An organisation at the 40th percentile of a wide distribution is not meaningfully different from one at the 60th.
  6. Is the benchmark stable year-to-year in method? A changing sample or definition makes year-over-year movement in the benchmark uninterpretable.

Using benchmarks defensibly

Where a benchmark passes the filter, its proper role is as a range, not a target. Plans should state the source, the definition alignment (or misalignment), and the percentile band the organisation sits in. Where definitions differ, the internal metric should be recomputed under the benchmark’s definition before comparison, rather than the comparison being made across definitions.

Often the more defensible reference point is the organisation’s own history, segmented properly. Internal trend under a frozen definition carries no sample-construction uncertainty at all.

What we cannot claim

None of this implies vendor benchmarks are produced in bad faith. Most are honest aggregations of available data, and the limitations follow from the economics of collecting it. Our claim is narrower: a benchmark is a measurement with a design, and an undisclosed design makes the measurement unusable for decisions with consequences. The data always wins over the narrative — but a benchmark without its methodology is mostly narrative.

Cite Workforce Data Lab, research desk. “How Vendor Benchmark Reports Get Constructed — and a Six-Question Filter for Using One as a Planning Input.” workforcedatalab.com, 13 January 2026. https://workforcedatalab.com/posts/2026-01-13-how-vendor-benchmark-reports-get-constructed-and-a-six-question-filter/

Further reading

from the same desk