AI Detector False Positives: The Legal Risk of Acting on a Score
A manager pastes a report into a detector, gets back "92% AI-generated," and starts a conversation that ends in a termination. Six months later the question in front of a hearing officer is not whether the employee used AI. It is what the 92% measured, who can explain it, and why the company treated a probability about the shape of sentences as proof of an act.
What the Number Is Measuring
These tools do not recognise text. They estimate how surprising a passage is to a language model — roughly, how often the next word is the word a model would have picked, and how much that varies across the passage. Machine-generated prose tends to be smoother and less surprising than human prose, so smoothness becomes the proxy for machine authorship.
Everything that makes writing conventional therefore raises the score. Writing in a second language. Writing to a template. Technical documentation. Legal and medical prose. Anyone who ran their draft through a grammar assistant, which is most people. The score is a measurement of style, and style is not conduct.
Four Different Claims, One Bad Decision
- •Disparate impact if false positives track national origin or disability
- •Wrongful termination where the stated cause cannot be evidenced
- •Unemployment benefit disputes the employer loses on proof
- •Retaliation claims when detection is applied selectively after a complaint
- •Union and contract grievances over just-cause standards
- •Defamation risk in repeating the accusation to colleagues or references
- •Consumer-reporting duties when an outside vendor scores a candidate
- •State AI and automated-decision notice rules triggered by screening use
- •Confidentiality breach from pasting client work into a third-party tool
- •Contract termination disputes with freelancers and agencies
The Disparate-Impact Problem, Precisely
A policy does not have to be aimed at anyone to be unlawful. If a neutral screening practice produces meaningfully worse outcomes for a protected group, the employer carries the burden of showing it is job-related and consistent with business necessity — and that a less discriminatory alternative was not available. A detector whose errors cluster on second-language writers fails every part of that test at once: the connection to job performance is speculative, the vendor cannot explain any individual score, and an obvious less-discriminatory alternative exists, which is asking the person for their drafts.
The Accused Cannot Win, Which Is the Tell
Ask what evidence would clear someone. Version history can be manufactured. Drafts can be fabricated. A confident denial reads as a confident denial whether or not it is true. When a process offers the accused no way to establish innocence, it is not an investigation — it is an accusation with paperwork, and that is the characteristic every adjudicator is trained to notice. The same structural defect is why academic institutions that adopted detectors early have been walking the practice back.
A Process That Survives Review
Before Any Detector Is Used
- ☐Publish an AI use policy: permitted, must-disclose, prohibited
- ☐Decide in advance that a score can start an inquiry and cannot end one
- ☐Check whether pasting the work into a vendor breaches a client confidentiality term
- ☐Confirm whether third-party scoring pulls you into consumer-reporting duties
- ☐Apply any detection uniformly, not just to people already under scrutiny
If You Suspect Undisclosed Use
- ☐Ask in writing for process, drafts, sources and version history
- ☐Compare against known samples of the person's work
- ☐Document the conduct you found, not the percentage you were shown
- ☐Give a real opportunity to respond before any adverse step
- ☐State the cause in terms you could prove without the tool
Frequently Asked Questions
Are AI detectors accurate enough to rely on?
Not for an adverse action against an individual. Even a tool with a genuinely low error rate across a large corpus produces a meaningful number of wrong answers at the level of one person's document, and it cannot tell you which answer is the wrong one. Accuracy claims are aggregate statements; a termination is a single case.
Why do second-language writers get flagged more?
Detectors score predictability, and writing produced in a second language is often more conventional in vocabulary and sentence construction than native prose. Translation and grammar tools smooth it further. The result is a higher machine-likeness score for reasons that have nothing to do with how the text was produced.
Can we be sued for accusing someone wrongly?
Repeating a false accusation of dishonesty to people who do not need to know it is the core of a defamation claim, and internal circulation is not automatically privileged. Separately, an adverse action you cannot evidence supports wrongful-termination and unemployment claims, and a screening practice with skewed error rates supports a discrimination claim.
Does the FCRA apply to running a candidate's writing sample through a tool?
It can. If an outside vendor assembles or evaluates information about a person and you use the result in an employment decision, consumer-reporting obligations may attach — disclosure, authorization, pre-adverse-action notice with a copy of the report, and a chance to dispute. Many teams skip these because a detector does not feel like a background check.
What about freelancers and agencies rather than employees?
The employment statutes may not reach the relationship, but the contract does. Terminating for cause on an unprovable allegation invites a breach claim and a fight over unpaid work, and if you told anyone else why the engagement ended, the defamation exposure is identical.
What is the actual fix?
A written policy plus a process requirement. State where AI is allowed and where it must be disclosed, and require that any adverse action rest on documented conduct — an undisclosed use the person acknowledges, a fabricated citation, a verifiably copied passage — rather than on a score. That converts an unfalsifiable accusation into a checkable fact.
Related Reading
The Same Question, Asked of Your Website
The failure here is acting on a number nobody can trace back to a specific, nameable defect. That is the difference between a score and a finding, and it is worth applying to every automated report your team acts on.
A scan of your own site returns the rule, the element and the line rather than a percentage. Run a free scan and see what a finding looks like when it names the thing at fault.