Workforce Data Lab
people-analytics Workforce Data Lab · research desk

Skills Data Quality: Inferred vs Attested Skills, and How to Decide with Bad Mirrors

Every skills-based initiative — talent marketplaces, workforce planning, succession analytics — inherits the error structure of whatever produced the skills data. That provenance is almost always one of two processes: inference (resume and ATS parsing, profile extraction) or attestation (self-report, manager endorsement). These two sources fail differently, predictably, and in ways that matter for which decisions each can support. This piece maps the error profiles and proposes a decision-bounding framework for the common situation: no ground truth available.

The two pipelines and where they break

Inferred skills are extracted from documents: resumes, job histories, project descriptions. Their error profile has a recognisable shape:

  • Staleness by construction. A resume records what a person once chose to advertise. Skills decay silently; nothing in the parsing pipeline deletes a competency last exercised nine years ago. Inferred data systematically overstates current capability.
  • Context bleed. Parsers cannot reliably distinguish “managed a team using Kubernetes” from “uses Kubernetes.” Proximity to a term becomes a skill claim. The false-positive direction is consistent: inference inflates.
  • Title-driven projection. Job titles carry strong priors, and extraction models lean on them. Two people with identical documents modulo title receive different inferred profiles. This imports every historical inequity encoded in who got which title.
  • Taxonomy collapse. Synonyms, versions, and adjacent tools must be normalised into a controlled vocabulary; every collapse decision destroys information, and the error compounds at reporting time.

Attested skills are declared by a person or their manager. Their failure modes are behavioural rather than mechanical:

  • Strategic inflation. When employees learn that skills profiles feed staffing, promotion, or project allocation, self-report becomes an instrument of career advancement. The distortion is rational and directional.
  • Self-assessment error. Self-ratings of proficiency are poorly calibrated in both directions — documented in the assessment literature as a general property of self-reported competence, not a quirk of any population.
  • Manager-sparsity and recency. Managers attest to what they have recently observed. Portfolio breadth is underreported; yesterday’s project is overreported. Coverage is thin anywhere a manager’s line of sight does not reach.
  • Polarised participation. Attestation systems get high coverage among the motivated and the managed, and near-zero coverage among the disengaged — which is to say, precisely where an attrition or capability risk might be developing.

Neither error profile is a reason to discard the source. They are reasons to stop treating either as the skills record.

Concordance as the only free lunch

Where inferred and attested data agree, confidence is earned cheaply. Where they disagree, the disagreement itself is informative, and the direction of disagreement is diagnostic: inferred-but-unattested skills are likely stale or context-bleed artefacts; attested-but-uninferred skills are likely recent, tactical, or aspirational. An illustrative framing — constructed, not measured: if the two pipelines agree on a skill for a given employee, the probability the claim is real is materially higher than the base rate of either source alone; the operational value is in the disagreement queue, which is a finite, reviewable list.

This is why we recommend every skills record carry explicit provenance fields at the claim level, not the profile level: source = inferred | self_attested | manager_attested | assessed, asserted_at, last_verified_at. Aggregates computed without these fields are unverifiable by construction.

Bounding decisions without ground truth

Ground truth — a validated assessment of who can actually do what — almost never exists at scale. The alternative is not to abandon decisions but to bound them. Four practices:

  1. Decision sensitivity analysis. For each decision the data will support, ask: how wrong does the input need to be before the preferred option changes? If a staffing decision flips only if 40%+ of claimed skills are false, modest data error is tolerable. If it flips at 10% error, the decision requires verification cost the organisation is probably not paying. State the break-even error rate next to the decision, in the same document.
  2. Calibration sampling. Audit a random sample of claims with human review — a structured conversation or work-sample check. The point is not precision; it is to put an upper bound on the error rate with a defensible method, so sensitivity analysis runs on evidence rather than vibes.
  3. Match decision tier to data tier. Inferred and attested data can support aggregate questions — capability distribution, gap surfaces at department level, trend detection — where individual errors wash out. They should not support individual consequences — promotion denials, performance processes, redundancy selection — without a verified claim. This is a governance line, and it should be documented as one.
  4. Degrade gracefully. When provenance is unknown or verification is older than its shelf life, the correct representation is lower confidence, not absence. Squashing uncertain data to zero manufactures false negatives; leaving it unweighted manufactures false positives. Keep the field, discount the weight.

What we cannot claim

We do not know, and cannot know without assessment data, the true error rates of either pipeline in any given organisation. The error-profile claims above are directional and documented in the assessment and information-extraction literatures; their magnitudes are organisation-specific. Anyone presenting a precise skills-coverage percentage without a calibration sample behind it is reporting a number whose error bars are unknown — which for skills data is worse than reporting none.

The skills record is a mirror of the workforce, and both available mirrors are warped in known directions. The defensible posture is not to pick the less warped mirror. It is to publish the warpage: provenance on every claim, agreement and disagreement surfaces made visible, sensitivity bounds on every decision, and a standing calibration sample so the error rate is estimated rather than assumed. The data always wins over the narrative — provided the data admits, on its face, where it came from.

Cite Workforce Data Lab, research desk. “Skills Data Quality: Inferred vs Attested Skills, and How to Decide with Bad Mirrors.” workforcedatalab.com, 02 December 2025. https://workforcedatalab.com/posts/2025-12-02-skills-data-quality-inferred-vs-attested-skills-and-how-to-decide-with/

Further reading

from the same desk