Workforce Data Lab
people-analytics Workforce Data Lab · research desk

GenAI in People Analytics: Where LLMs Help, Where They Fabricate, and a Validation Protocol

Large language models have entered the people-analytics workflow faster than the validation practices around them. Eighteen months of watching teams deploy them has produced a clear pattern: the failures concentrate in specific task types, not in “AI usage” generally. This piece separates where LLMs genuinely add capability from where they manufacture confident error, and closes with the validation protocol we now consider the floor for AI-assisted people analytics.

Where LLMs genuinely help

The reliable use cases share a structure: the model transforms text into structured, checkable categories, and a human can verify any given output.

Open-text theme extraction. Engagement-survey comments, exit-interview notes, and free-text fields have historically been underused because manual coding does not scale. LLMs handle the volume. The defensible pattern is a fixed codebook, the model as coder, and a human-coded audit sample against which the model’s agreement rate is measured — a quality gate, not a one-off check. Theme frequency at corpus scale is a legitimate LLM output; theme existence in any individual comment is checkable by reading the comment, which keeps the task honest.

Qualitative coding at scale with drift-resistant definitions. A codebook plus a prompt plus a frozen model version produces apply-the-same-rule-tomorrow consistency that human coding teams — subject to fatigue, turnover, and definitional drift — struggle to match. Consistency is not accuracy, but for longitudinal text analysis, a consistent coder with a measured error rate beats an inconsistent one with an unmeasured rate.

First-pass narrative over verified numbers. Asking a model to write commentary given a computed table — percentages supplied in the prompt, generated by code, not by the model — is safe, provided the drafting prompt contains the verified values and the output is checked against them. This is the one generative use we endorse broadly, because every claim in the output is traceable to a supplied number.

Where they fabricate

The failure modes are equally specific.

Arithmetic and aggregation. An LLM asked to compute statistics over raw data — “what share of comments mention pay?” — will produce a plausible number that is not a count of anything. It is generating text shaped like an answer. The rule is absolute: no statistic in a published analysis may originate inside the model. Counts, percentages, differences, and trends are computed in code; the model may only narrate values that were computed elsewhere.

Citation and provenance. Asked why a theme matters or what prior work found, models invent studies, dates, and sample sizes — fluent, formatted, and false. Any reference an LLM produces must be treated as a search query, not a citation, until a human has located the actual source.

Confident theme invention under thin data. At small corpora, models resolve ambiguity in favour of coherence. A handful of vivid comments becomes a stated theme at a stated prevalence. The fabrication is not malicious; it is the model completing a pattern. Thin-data outputs need heavier human review, not lighter, because the prose quality is inversely related to the evidence.

Consistency decay across versions. A theme taxonomy applied under one model version is not guaranteed identical under the next. Any longitudinal text programme that cannot pin its model version is re-baselining its instrument silently — the same wording-changes-break-trends problem as survey instruments, at higher velocity.

The validation protocol

Six practices, all cheap relative to the cost of a fabricated number reaching an executive deck:

  1. Numbers from code, prose from models — never the reverse. The pipeline computes every statistic deterministically. The model receives computed values as input. The output is scanned to confirm every number it contains matches a supplied value; any unmatched number fails the run.
  2. Audit-sample coding with a measured agreement rate. For classification tasks, human-code a random sample each cycle and report the model’s agreement against it. The agreement rate travels with the deliverable. When it degrades — new language, new topics — re-specify the codebook; do not average the degradation into the reporting.
  3. Pin model version and prompt; log both. Reproducibility requires that last quarter’s analysis be re-runnable. Version, prompt text, and codebook hash are part of the analysis artefact, like a survey instrument is part of a survey report.
  4. Citation quarantine. Nothing an LLM names — study, dataset, benchmark — enters a deliverable without independent verification. One sentence of protocol; it eliminates an entire failure class.
  5. Pre-registered themes for trend claims. If a theme’s prevalence will be trended over time, the theme definition is frozen in advance. Emergent-theme discovery is exploratory output and is labelled as such, never trended.
  6. Privacy floor on inputs. Employee comments enter the model only under the organisation’s data-processing posture — retention, residency, and use restrictions stated, same as any processor. A validation protocol that excludes the data-handling question is not a protocol, it is a style guide.

What we cannot claim

We make no claim about which models perform best at these tasks, and we are deliberately avoiding benchmark citations — the evaluation literature is moving faster than any recommendation would survive. The split above is structural, not model-dependent: transformation-and-classification tasks admit verification; generation-over-raw-data tasks do not. That distinction will outlast any leaderboard.

GenAI earns its place in the analytics stack the same way every other instrument does: by accepting measurement. Pinned, audited, and quarantined from arithmetic, it is the best tool the field has ever had for qualitative scale. Unmeasured, it is a confident colleague who never shows work. The data always wins over the narrative — and the first job of the protocol is making sure the data wasn’t written by the narrator.

Cite Workforce Data Lab, research desk. “GenAI in People Analytics: Where LLMs Help, Where They Fabricate, and a Validation Protocol.” workforcedatalab.com, 19 May 2026. https://workforcedatalab.com/posts/2026-05-19-genai-in-people-analytics-where-llms-help-where-they-fabricate-and-a-v/

Further reading

from the same desk