The Vendor Says the Model Is Validated. Validated Against What Job?
Validity is not a property a model carries around with it. It is a claim about a specific inference, for a specific job, at a specific cut score — and the part that makes it true is the part no vendor can sell you.
The document most employers have is the wrong one. A bias audit reports who gets screened out. The question a disparate impact claim asks is whether the thing being measured predicts performance on the job in question. Those are answered by different studies, and the second one requires a job analysis that only the employer can produce.
Four Things "Validated" Is Routinely Taken to Mean
Each of these appears in procurement files as though it closed the question. None of them does.
A bias audit is not a validation study
A bias audit reports selection rates by group. It answers whether the tool screens out protected groups at different rates. It says nothing about whether the thing being measured predicts performance on your job. An employer facing a disparate impact claim needs the second document, not the first, and most procurement files contain only the first because that is what a city ordinance made mandatory.
'Validated' is not a property of a model
Validity is a property of an inference drawn from a score, for a defined job, in a defined population, at a defined cut. The same model can be a valid predictor for a high-volume contact-centre role and meaningless for a field technician. A vendor sentence that says the assessment is validated, with no job named, is not a claim that can be true or false.
Accuracy against a hiring-manager rating is not job relatedness
Many vendor studies report how well the model reproduces past human decisions. If those decisions carried the skew being challenged, a model that reproduces them faithfully is a high-accuracy model and a weak defense. The criterion has to be a measure of job performance, not a measure of who got picked before.
A technical manual is not an evidence file
Vendors ship a technical manual describing the construct and the psychometrics. What an employer needs is a record showing this employer analysed this job, chose this tool for reasons tied to that analysis, set this cut score on a stated basis, and monitored the result. That record is made by the employer and cannot be bought.
Five Kinds of Evidence, Five Different Failure Modes
Criterion-related validity
What it shows: Scores correlate with a measure of performance on the job — supervisor ratings, objective output, retention, error rates.
Where it breaks: The correlation is only as good as the performance measure behind it. Supervisor ratings carry their own skew, and a study that validates against them inherits it. Ask what the criterion was, who produced it, and whether the criterion itself was checked for group differences.
Content validity
What it shows: The assessment samples the actual work — a coding exercise for a developer, a call simulation for support.
Where it breaks: Content validity is the friendliest route for a small employer because it needs no sample size, but it only supports tasks that are directly observable work behaviours. It does not support inferred traits. A personality inventory or a video-derived 'engagement' score cannot be defended this way, and that is exactly where AI screening tends to sit.
Construct validity
What it shows: The assessment measures a trait, and the trait matters for the job.
Where it breaks: The most demanding route and the one most often asserted without support. It needs evidence for the construct, evidence that the instrument measures it, and evidence linking it to job performance. Three links, three places to break.
Transportability
What it shows: Borrowing a validity study conducted elsewhere, on the grounds that the job is substantially the same.
Where it breaks: This is the route nearly every AI vendor is implicitly relying on, and it carries conditions: the original study has to be technically sound, the jobs have to be shown comparable by job analysis, and the original context must not differ in ways that matter. The comparison has to be documented by you. 'Our model was validated across 4 million assessments' is not a transportability argument — it names no job.
Differential validity and fairness
What it shows: Whether the score predicts performance equally well across groups, and whether it systematically under- or over-predicts for any of them.
Where it breaks: Distinct from the adverse impact ratio and usually missing. A tool can pass a four-fifths check and still predict badly for one group; it can also fail the ratio and still be defensible if the validity evidence is strong and no less discriminatory alternative exists. These are different questions and need separate answers.
The Job Analysis Is the Load-Bearing Document
Every route to validity runs through a description of the work that was produced by method rather than by memory. Four questions establish whether yours exists.
Who wrote the job description the vendor mapped to?
Usually a recruiter, at speed, reusing the last posting. A job analysis is a different artefact: observation or structured interview with incumbents and supervisors, a list of tasks with frequency and criticality ratings, and the knowledge, skills and abilities each task requires. If the tool was configured against a job posting rather than a job analysis, the link between the assessment and the work is an assumption.
Does the analysis distinguish entry requirements from things learned on the job?
It has to. A requirement that the successful applicant already possesses something the employer teaches in the first fortnight is hard to defend when it screens out a protected group, and it is the most common finding when a real analysis is finally done. The analysis is also what identifies the essential functions that an accommodation conversation later turns on.
Was it redone when the job changed?
Roles drift, the model does not notice, and the file goes stale quietly. Date the analysis and re-examine it when the role materially changes or on a fixed cycle. A five-year-old analysis for a role that was remote-first two reorganisations ago will not support the tool sitting in front of it today.
Does one analysis cover every requisition the tool touches?
Almost never, and this is where scope quietly outruns evidence. Tools get switched on for one role, work acceptably, and spread across a job family nobody re-analysed. The defensible unit is the job or the demonstrably similar job group, not the applicant tracking system.
The Cut Score Is a Separate Decision and It Is Usually Undocumented
Write down what the score is for
Rank ordering, a pass threshold, or a band feeding a human decision. Each implies different evidence. A threshold that rejects is the strongest claim and needs the strongest support.
Justify the number, not just the instrument
A cut score is a separate decision from the assessment, and it is the decision that produces the impact. A defensible one is set on a reasoned basis — expected performance at the threshold, the level of proficiency the work genuinely requires, or a normative argument — and recorded at the time. A number inherited from a vendor default has no such record.
Check impact at the number you actually use
Impact is a function of the threshold. A tool audited at one cut and deployed at another has been audited for a decision nobody makes. Recompute selection rates at the live setting, including any downstream ranking or auto-advance.
Look for the less discriminatory alternative before you are asked
Once impact is shown and job relatedness is established, the remaining question is whether an equally effective alternative with less impact exists. Testing a lower cut, a different weighting or a structured human step and recording the result turns that question from a surprise into a document.
Keep the applicant-flow data that lets anyone recompute it
Records of applications, dispositions and assessment scores with the demographic data needed to compute rates. Without the flow data no rate can be computed after the fact — including by you, when deciding whether a claim has anything behind it.
Six Questions That Belong in the Contract, Not the Demo
Name the job families the validity evidence covers
A vendor that can answer this specifically is a vendor whose study exists. A vendor that answers with a volume of assessments is telling you it does not transport.
Produce the criterion definition
What performance measure the scores were validated against, who generated it, and whether it was examined for group differences. This single answer separates a study from a marketing claim.
State the sample: size, roles, employers, dates
Small samples, single-employer samples and samples predating a model change all limit what the study supports. Dates matter most — a study run against a previous model version is evidence about software that is no longer in the loop.
Disclose every model change that affects scoring
Contract for it. A silent retrain or feature change resets the evidence and nobody tells the buyer. Require notice, a re-run of the impact figures, and the right to pause the tool pending review.
Get the fairness analysis, not just the ratio
Selection rates by group, plus differential prediction if the sample supports it. Ask which groups were too small to analyse; that answer is usually where the real exposure is.
Secure your right to the underlying data
Scores, features and dispositions for your own applicants, in a portable format, during the term and after it. When a claim arrives two years later, a vendor relationship that has ended should not take the evidence with it.
Questions Employers Ask
Our vendor gave us a bias audit. Is that enough?
It satisfies the ordinance that demanded it and it does not answer the question a discrimination claim asks. A bias audit reports selection rates by group at a point in time. The defense to an established disparate impact is that the practice is job related for the position in question and consistent with business necessity, and that is shown with validity evidence tied to a job analysis — a different document with a different method. Treat the audit as a monitoring instrument that tells you whether you have a problem, and the validation file as the thing that answers for it. Employers who have only the audit typically discover the gap at the point where they are asked to produce the second document under a deadline.
Can we rely on the vendor's study instead of running our own?
Sometimes, through transportability, but it is a conditional route rather than a free one. Borrowing an external study requires that the original study be technically sound, that your job be shown comparable to the job studied through job analysis, and that contextual differences not undermine the inference. Each of those is an argument you have to make and document — the vendor cannot make the second one for you because it does not know your job. In practice the work is a job analysis plus a written comparison against the study's job description, which is a contained piece of work for one role and a substantial one across a whole requisition catalogue. Do it for the roles with volume first, because that is where impact is computable and where a claim is most likely to originate.
We are a small employer with no statistical sample. What can we defend?
Content validity is usually the realistic path, and it constrains the tool rather than the paperwork. A work sample that mirrors an actual task — a real support ticket, a real spreadsheet, a real code change — can be supported by a careful job analysis and subject-matter expert judgement without a criterion study. What that route will not support is an inferred trait: a personality score, a video-derived engagement measure, a game-based construct. If the tool you are buying measures something you cannot point to in the work, a small employer cannot defend it on content grounds, and there is no sample to defend it on criterion grounds either. The procurement decision and the legal position are the same decision.
Does a human reviewing the output fix the problem?
Only if the human can actually change the outcome, and only for the stages the human sees. A reviewer looking at a shortlist the model produced cannot restore a candidate the model already removed, so screening-out steps carry their exposure regardless of what happens downstream. Meaningful review means the reviewer sees the candidates who were filtered, has the information needed to disagree, has time proportionate to the decision, and disagrees sometimes — a review step with a recorded override rate of zero across thousands of decisions is evidence that it is not functioning as a control. Record overrides. The rate is both your operating signal and your evidence that the safeguard is real.
How long do we have to keep the records?
Longer than most applicant tracking configurations default to, and the retention has to cover three separate things: the applicant flow data needed to compute selection rates, the assessment scores and features for individual candidates, and the validation and job analysis file itself. Federal recordkeeping duties set a floor that lengthens once a charge is filed, and a preservation obligation attaches when litigation is reasonably anticipated — which means the moment a charge notice arrives, deletion routines have to stop. The common failure is automated: a data minimisation policy written for privacy compliance quietly destroys the evidence needed for an employment defense. Reconcile the two schedules deliberately and write down which one wins for applicant data.
The One-Folder Test
Pick the requisition your screening tool touched most times last quarter. Try to assemble, from what already exists, a single folder containing: the job analysis, the validity evidence naming that job, the cut score and the reason for it, the selection rates at that cut, and the applicant flow data behind them.
Whatever is missing from that folder is not a paperwork gap. It is the part of the defense that would have to be constructed after the fact, by people who were not there, under a deadline set by somebody else.
Related Reading
- The four-fifths rule and AI screening — how the impact ratio is computed, and what it does not settle.
- Hiring records retention — what has to survive long enough to be recomputed.
- NYC, Illinois and Colorado compared — which jurisdictions demand an audit, and what each one counts.