A Model Decided the Claimant Could Work. Now Explain It.
Disability claims are the rare decision where the protected characteristic is the subject matter. That makes the usual defence — we never collect that attribute — unavailable, and it makes a model trained on a decade of prior denials a very efficient way to repeat them.
The Decision Nobody Recorded as a Decision
Ask a benefits operation where AI touches a claim and the answer is usually "nowhere that matters — a person decides." Follow the claim instead. It arrives, is scored, and on the strength of that score is routed to a fast-track queue or an investigative one. The investigative queue orders an independent medical examination, a vocational assessment, sometimes surveillance. The fast-track queue approves. Two claimants with similar files receive materially different processes, and the fork was chosen by a model that never produced a denial and never appears in a denial letter.
Every regime that governs this work cares about that fork. Discrimination law cares because scrutiny is a burden distributed unevenly. Claims-handling regulation cares because it prescribes how claims are investigated, not only how they are decided. And a reviewing court asked whether the process was full and fair cares a great deal about a step in the process the claimant was never told existed.
Why the Usual Bias Defence Does Not Work Here
In hiring or lending, teams defend a model by showing it never sees race, sex or disability. In disability claims the model is reasoning about the medical condition directly — that is its job. So the question moves from which attributes the model can see to which conditions it treats worse, and the honest answer usually falls out of the training data without anyone intending it.
Historical claim files are a record of past decisions, not of true functional capacity. Where those files reflect a long institutional scepticism toward conditions without a clean objective marker, a model fitted to them learns the scepticism as signal. It will predict denial for those claims accurately, because the organisation did deny them — and it will be measured as accurate by the very outcomes it is reproducing. This is the central measurement trap of the category: label leakage from the decision you are trying to audit.
Eight Questions a Fiduciary Should Be Able to Answer
- Which models and guideline versions touched this claim, on what dates?
- What did each output, and what did the reviewer do with it?
- What was the stated reason for the adverse determination, and is it the operative one?
- Which internal rules or criteria were relied upon, and would we produce them on request?
- How do referral, surveillance and denial rates vary by condition category?
- How do they vary by claimant age and sex within condition category?
- What is the override rate, and is it high enough to call the human review real?
- Can we reproduce the score for this claim today from the retained record?
The override rate is the most diagnostic item on that list. A reviewer who accepts the model's recommendation in almost every case is not providing oversight, whatever the procedure document says, and the figure is one query away in any system with an audit trail. Organisations that measure it are frequently surprised, in both directions.
Surveillance Triage Deserves Its Own Review
Of every downstream action a claims model can trigger, surveillance is the most intrusive and the least examined. It is typically governed by an internal referral policy rather than by anything a claimant sees, and the referral criteria are exactly the sort of internal guideline a full-and-fair-review request reaches. When the referral is model-driven, two additional problems attach: the distribution of surveillance across condition categories becomes a measurable disparity, and the model's inputs may include social-media or lifestyle signals whose collection has its own privacy consequences.
Treat a surveillance referral as a consequential decision in its own right — with a documented reason, human authorisation that is not a rubber stamp, and periodic distribution testing — rather than as an investigative detail beneath the level of governance.
Questions Carriers, TPAs and Plan Sponsors Ask
We are the plan sponsor, not the carrier. Does any of this reach us?
It reaches you through the fiduciary duties attached to plan administration and through your selection and monitoring of the entity doing the work. Delegating claims administration does not delegate the duty to select and monitor prudently, and a monitoring process that never asks how claims are adjudicated is thin. The practical ask is small and almost never made: request the vendor's description of any automated tools used in adjudication, the governance around them, and any outcome testing by condition category. A vendor unable to answer is telling you something useful, and the request itself is evidence of monitoring if the question is ever litigated.
Does telling the claimant a model was used invite challenges?
The alternative invites worse ones. A claimant who discovers a determinative model input during discovery gets two arguments — the substance of the decision and the concealment — and the second tends to colour a reviewer's reading of the first. Disclosure also has a quieter benefit: the specific-reason requirement forces the letter to name what actually drove the outcome, which surfaces internally when the letter cannot honestly do so. Organisations that adopt plain-language explanations generally find appeal volume roughly stable and appeal quality higher, because claimants send evidence addressed to the real issue instead of guessing.
How do we test for disparity when we do not hold demographic data?
Start with what you certainly hold: condition category, occupation class, claim duration, referral and surveillance rates, approval and termination rates. Condition-level disparity is the legally salient axis here and needs no protected attribute at all. Age and sex are typically already in the file for underwriting or benefit-calculation reasons. For anything genuinely absent, a separated, access-controlled evaluation dataset used only for bias measurement is the pattern regulators and the newer AI statutes contemplate, and it is not the same thing as adding a demographic feature to the model.
The model is a vendor's and they call it proprietary. Now what?
Proprietary is a commercial position, not a legal exemption from the disclosure and governance duties that sit on you. Negotiate the needed rights before signing: documentation of the model's purpose and inputs, outcome testing by cohort, notification of version changes, retention of per-claim inputs and outputs in your record, and the right to disclose whatever a regulator or a claimant's representative is entitled to receive. Where a vendor refuses, the meaningful question is whether you can defend a decision you cannot explain — and the answer determines whether that product is usable in adjudication at all, however good its accuracy numbers look in the deck.
What is the single highest-value change to make first?
Write the model's contribution into the claim file at the moment it happens. Not a policy, not a committee — a field. Most of the exposure in this area is evidentiary rather than substantive: the decisions are often defensible and the files cannot show it, because the score lived in a pipeline with a ninety-day retention and nobody can now say why a claim was routed for investigation in the first place. A recorded model version, input snapshot, output and reviewer action costs one schema change and converts an unreconstructable decision into an ordinary defensible one.
Run the Override Rate This Week
One query: of the claims where a model recommended referral, denial or a duration, how often did the human reviewer reach a different answer? If the number is close to zero, the human-in-the-loop in your procedure document does not exist in your operation, and every defence built on it is built on the query you have not run.
Then break the outcome rates down by condition category. Those two numbers decide how much of the rest of this list you actually need.