RatedWithAI

RatedWithAI

Accessibility scanner

Algorithmic DiscriminationSeptember 8, 2026

AI Sourcing Built the Pool. Nobody Tested Who Never Got Into It.

Every adverse-impact number an employer produces starts at the application. Sourcing tools decide who ever becomes an applicant — and the people they skip are absent from the data set the audit is run on.

No denominator
Adverse-impact ratios need a pool; sourcing decides the pool before one exists
Not an applicant
Passive candidates usually fall outside the applicant definition — and inside the liability
Weeks of history
Vendor search logs roll over long before a charge or audit arrives

The Stage That Compliance Programs Skip

Hiring compliance has a well-worn shape. You define the requisition, you collect applications, you measure selection rates by stage, and you compare them. Every part of that machinery begins at the moment somebody applies, because that is the moment the employer traditionally acquired knowledge of the person.

AI sourcing inverts the order. The tool holds an index of tens of millions of profiles, infers skills and seniority from unstructured text, ranks them against a job description or against a model of who you hired before, and hands a recruiter a list of forty names. Nobody applied. Nothing entered the ATS. No selection rate moved. A selection decision was made all the same, and it was made on a population you cannot see in any report you currently produce.

Three product categories sit here, and they behave differently enough to be assessed separately: external sourcing platforms searching a third-party index, talent rediscovery tools ranking your own past applicants, and outbound matching engines that decide which candidates see which job messages. The last is closest to advertising delivery and inherits its problems; the middle one is the most common and the most overlooked.

Four Mechanisms That Produce Skew Without Anyone Choosing It

Training on your own hiring history

Rediscovery and match-scoring tools are commonly fitted to the profiles of people you previously hired or advanced, because that is the only labelled outcome data an employer has.

Why it produces exposure: Whatever composition your past hiring produced becomes the target the model optimizes toward. If a function was homogeneous for reasons having nothing to do with merit, the tool encodes that as the definition of a good match and reproduces it at speed, across every requisition, without a human ever articulating the criterion.

Inferred attributes standing in for protected ones

Profiles rarely state skills directly, so tools infer them from employer names, school names, job titles, gaps in dates, group memberships and writing style.

Why it produces exposure: Several of those inferences correlate tightly with age, national origin, disability and caregiving status. Graduation year proxies age. An employment gap proxies medical leave or caregiving. Institution names proxy national origin. None of it is stated, all of it is learned, and the ranking that results is difficult to explain after the fact.

Geographic and platform coverage

The index is not the labor market. It is whoever maintains a detailed public profile on the platforms the vendor scrapes or licenses.

Why it produces exposure: Profile completeness varies systematically by age, occupation, income and region. A search that surfaces only richly documented profiles is filtering on a variable that tracks protected characteristics, and it does so before your job description has any effect at all.

Engagement-optimized outreach

Where the tool decides who to message, or who sees a posting, the objective is usually predicted response or predicted hire — not coverage of the qualified population.

Why it produces exposure: Optimizing delivery for response rate reliably concentrates delivery on the demographic that responded most in the training window. This is the failure mode already litigated in ad delivery, and it does not become a different mechanism because the surface is a recruiting inbox.

Building a Denominator for a Stage Designed Without One

The analytical problem is real, not rhetorical: you cannot compute a selection rate without knowing who was available to be selected. Three workable reference populations, in descending order of defensibility:

  • The searched set. For a rediscovery tool ranking your own database, the reference population is the filtered database. This is the strongest comparison available because both sides of it are yours, and it isolates the tool's behavior from labor-market supply.
  • The vendor's candidate universe under your filters. For an external platform, ask the vendor for the size and aggregate composition of the result set your query matched before ranking. Some can produce it. A vendor that cannot describe what the model chose from has told you something worth writing down.
  • An external availability benchmark. Occupation and geography-specific labor-force estimates are the fallback where nothing internal exists. Weakest of the three, because it mixes the tool's behavior with the shape of the underlying market — but it is far better than comparing surfaced candidates to nothing.

Whichever you pick, run the comparison per requisition family rather than in aggregate. Aggregation across dissimilar roles is the standard way a real disparity in one function disappears into a portfolio-level ratio that looks acceptable.

The Sourcing-Stage Control Set

These are the controls that make the stage measurable. Most organizations have none of them, not because they were rejected but because the stage was never on the map.

1. Inventory and Scope
  • List every tool that ranks, scores or recommends candidates before an application exists
  • Include ATS features that surface past applicants — rediscovery is usually a module, not a purchase
  • Record for each whether it scores, or only filters on criteria a recruiter typed
  • Flag anything trained on your prior hiring outcomes for priority review
  • Confirm whether any of them fall inside a state or city bias-audit statute on function
2. Evidence and Retention
  • Export search parameters, filters and surfaced sets on a schedule into your own retention system
  • Capture the model version and configuration alongside each export
  • Retain the reference population definition, not only the results
  • Log who was contacted and the outcome, so the stage connects to applicant-flow data
  • Verify the vendor's default log retention window in writing — assume it is shorter than yours
3. Testing
  • Compare surfaced-set composition against the searched set, per requisition family
  • Re-run after any model version change, since prior results describe a prior model
  • Test whether inferred attributes (graduation year, gaps, institution) drive ranking
  • Check outreach delivery separately from ranking where the tool decides who gets messaged
  • Document the estimation method and its limits before the results exist, not after
4. Vendor Terms
  • Require disclosure of what population the model searches and how it was assembled
  • Require notice of model version changes, with prior-version availability during revalidation
  • Secure the right to export raw search and result logs in a machine-readable format
  • Ask what bias testing the vendor performed, on which population, and for which roles
  • Confirm cooperation obligations if you receive a charge or an audit request

Frequently Asked Questions

Our recruiters make the final call on who to contact. Does human review solve this?

Only if the review is capable of adding back what the ranking removed, and a recruiter looking at a list of forty names cannot restore the four hundred who were never listed. Human review is a meaningful control at the rejection stage, where the reviewer sees the candidate and the reason. At the sourcing stage the human sees only survivors, so the review validates the output of the model using the model's own output as the universe. If you want the human step to do real work here, give the recruiter the ability to see and search the unranked set, and measure how often they reach into it. If they never do in practice, the ranking is the decision regardless of who clicks send.

The vendor says their tool is bias-tested and certified. Is that enough?

It is a useful input and not a substitute for testing your deployment. Vendor testing evaluates the model on the vendor's population under the vendor's configuration for a generic role family. Your exposure comes from your filters, your job descriptions, your geography and — if the tool learns from your outcomes — your hiring history, none of which were in the vendor's test. The pattern regulators and plaintiffs have followed is to look at the deploying employer's results, because that is where the decision landed. Ask for the vendor's methodology and population, keep it in the file, and then run your own comparison on your own surfaced sets.

Is talent rediscovery riskier than external sourcing?

Usually yes, for a reason that is easy to miss: the population is your own past applicants, and the ranking model is often fitted to your own past hires. That closes a loop. Whatever composition your prior selection produced becomes the training signal for who gets resurfaced, and each cycle reinforces the last. It is also the tool most likely to be enabled as a feature rather than bought as a product, so it frequently never passes through procurement, legal review or the AI inventory. On the other hand, it is the easiest one to test properly, because you own both the reference population and the result set.

What about candidates who are outside our target geography or salary band? Is filtering them out a problem?

Filtering on genuine job-related requirements is ordinary and lawful; the issue is the criteria that ride along with them. Location filters carry demographic composition, and where the role can be performed remotely, a tight radius is hard to defend as job-related. Salary-band filters based on inferred current compensation raise their own problem, since inferred pay reflects prior pay and prior pay reflects historical disparity — several jurisdictions restrict compensation-history use for exactly that reason. The practical test is whether you can state the business necessity for each filter in a sentence, before you see the results.

We are a small employer without an analytics team. Where do we start?

Start with inventory and retention, which cost almost nothing and preserve the option to do everything else. Write down which tools rank candidates, turn on or schedule log exports, and keep the searched-set definition with each export. That alone moves you from having no record of a selection stage to having one. The testing can follow, and for a small hiring volume the honest answer is often that per-requisition statistics will be too sparse to be meaningful — in which case the useful work is qualitative: read the filters, read what the model was trained on, and remove criteria you cannot justify. Small volume is a reason to reason about the mechanism rather than a reason to skip the stage.

The Cheapest Fix Is the Export Schedule

Every other control in this article depends on having a record of what the tool did. Vendor search history rolls over on a product schedule that has nothing to do with your obligations, and once it is gone the stage is unauditable — by you, and by anyone asking you to explain it.

Set the export up first, then decide what to measure. The order matters, because the data you did not keep is the only kind you cannot go back and analyze.