RatedWithAI

RatedWithAI

Accessibility scanner

Algorithmic DiscriminationAugust 29, 2026

The Model Scored the Prose. The Rubric Asked About the Argument.

Automated scoring is sold as a workload fix and lands as a consequential decision about a student. The gap between what the rubric measures and what the model rewards is where the civil-rights exposure sits — and it falls hardest on the students institutions are specifically obliged to protect.

Title VI
National-origin disparate impact where English learners score lower
504 / IDEA
Accommodations must survive the scoring step, not just the task
Construct Drift
The model rewards fluency; the rubric claims to measure reasoning

A Grade Is a Consequential Decision

Algorithmic-accountability law has converged on a category it calls consequential decisions — the ones that determine access to employment, credit, housing, healthcare and education. Admissions is obviously in that category and gets the attention. Grading sits inside it too, and gets far less.

A course grade determines progression, placement, eligibility for programmes, financial aid standing and transcript record. A scored placement test determines which track a student enters. These are not soft outcomes, and an institution that automated them has automated a decision that its civil-rights obligations already covered.

The practical consequence: the analysis you would run on an AI hiring tool applies here nearly line for line. Validity for the stated construct, outcome testing across protected groups, accommodation, notice, appeal, records.

Construct Drift: What the Model Is Actually Scoring

The core technical problem has a name in assessment research long predating current models. A scoring system is valid only for the construct it was validated against, and automated scorers reliably learn proxies for human ratings rather than the reasoning the rubric describes.

  • Surface fluency features — length, lexical variety, idiomatic phrasing, conventional syntax — carry most of the predictive weight because they correlated with human scores in training
  • Students writing in a second language produce substantively sound arguments in linguistically atypical prose, and score low on features that are not the rubric's subject
  • Students using speech-to-text, or writing with a disability affecting fluency or mechanics, hit the same penalty through a different route
  • Dialect and register variation get scored as error rather than as difference, which is a national-origin and race question, not a style one
  • Model outputs are highly gameable in the opposite direction, so the tool rewards students coached on what it likes — an equity gap on top of the first one

This is why "the vendor says it correlates with human raters at 0.8" is not a compliance answer. Agreement with human scores says the model reproduces the rating pattern, including whatever disparities that pattern already contained.

The Legal Frames That Apply

Title VI — National Origin
Covers recipients of federal financial assistance. A scoring tool producing systematically lower outcomes for English learners raises a disparate-impact question, and language-access obligations run alongside it where instruction and appeal processes are English-only.
Section 504 and the ADA
Accommodation duties attach to the assessment as a whole. If the scoring step penalises the accommodated form of a student's work, the accommodation was not effective — and an inaccessible appeal or score-review interface is a separate violation.
IDEA
For students with IEPs, assessment methods and any changes to how progress is measured belong in the team's decision, not in a procurement decision made at district level and applied silently across classrooms.
State ADMT and AI Statutes
A growing set of state rules on automated decision-making technology reach education uses, adding notice, opt-out or human-review requirements and, in several, a documented risk assessment before deployment.
FERPA
Student work sent to a scoring vendor is an education record. The school-official exception has conditions — direct control over the data, use limited to the contracted purpose, and no repurposing of submissions for model training absent proper authority.

Compliance Checklist for Districts, Universities and Vendors

1. Validate Against the Construct, Not Against Human Raters
  • Require evidence the tool was validated for your student population and the specific task type, not a general writing benchmark
  • Treat rater-agreement statistics as insufficient on their own — they inherit any disparity present in the human scores used as ground truth
  • Ask the vendor which features drive the score, and reject 'proprietary' as an answer for a tool making consequential decisions about students
2. Run the Outcome Test You Would Run on a Hiring Tool
  • Compare score distributions across English-learner status, disability status, race and national origin — the disparity is measurable in data you already hold
  • Re-run it after every model update, because a vendor version change can shift the distribution without any notice to you
  • Document remediation when a gap appears; a found-and-addressed disparity is a materially different position from one discovered by a complainant
3. Keep the Human Decision Real
  • Use automated scores as one input to an educator's judgment rather than as the recorded grade, and record the educator's independent basis
  • Track the override rate — if instructors almost never depart from the machine score, the human step will be characterised as nominal
  • Require independent human re-scoring on appeal, blind to the original machine score, rather than a review of it
4. Notice, Accommodation and Records
  • Tell students and families in plain language when automated scoring is used and how to request review — silent deployment is the fact that makes complaints sympathetic
  • Verify accommodations survive the scoring step, including dictated, extended-time and alternate-format responses
  • Retain the submitted artefact, the machine output and the educator's decision together for the length of the grievance window and any applicable records-retention rule
  • Confirm in the vendor contract that submissions are not used for model training and that FERPA school-official conditions are met

Frequently Asked Questions

We only use AI scoring for formative feedback, not for recorded grades. Are we exposed?

Much less, and the distinction is worth preserving deliberately. It erodes when formative scores feed placement, intervention assignment or a participation component, or when instructors adopt the suggested score as the grade. If the output influences a consequential outcome in practice, the formative label will not carry the analysis.

Is an AI writing-detector flag the same issue?

Related but distinct. Detectors are accusation tools with documented false-positive skew against non-native writers, and the failure mode is an academic-integrity penalty imposed on unreliable evidence. The scoring analysis here is about grades produced without an accusation; a detector case adds due-process and defamation-adjacent dimensions on top.

The vendor's contract says their tool is bias-tested. Is that enough?

It is a starting point and not a defence. The obligation to students runs to the institution, and a general fairness assertion about a national dataset says little about your population. Ask for the disaggregated results, then verify against your own outcomes after deployment.

How do we handle a student who says the AI graded them unfairly?

Give an independent human re-score that does not see the machine output first, preserve the artefact and the original score, and log the outcome. A pattern of successful appeals concentrated in one student group is the earliest signal you will get that the tool has a disparity problem.

Does this apply to private schools and non-federally-funded institutions?

Title VI turns on federal financial assistance, so coverage varies, but ADA public-accommodation obligations, state civil-rights and consumer-protection law, state ADMT rules and contractual promises to families all still apply. Very few institutions are outside all of these at once.

You Already Have the Data to Test This

Unlike a hiring tool, where applicant demographics are often unavailable, a school already holds English-learner status, disability status and demographics alongside every score. The disparity analysis is a query, not a research project.

That cuts both ways. It means the test is cheap to run — and it means a complainant's expert can run it too, on data the institution is obliged to have kept.

Related Reading