RatedWithAI

RatedWithAI

Accessibility scanner

Biometric PrivacyAugust 29, 2026

The Camera Is Not the Risk. The Corpus Behind the Model Is.

Biometric statutes regulate obtaining, storing, disclosing and profiting from identifiers — verbs that reach a training set as squarely as a turnstile. If your product's face or voice capability came from someone else's scraped corpus, you bought an exposure that no amount of deployment-side consent repairs.

Obtain / store
Statutory verbs that attach at training time, long before any user sees the feature
Not consent
Public availability of a photograph does not authorise generating a template from it
Model deletion
A remedy regulators have pursued: destroy the data and what was built from it

Why Training Sits Inside the Statute

Read the verbs. Illinois BIPA speaks about collecting, capturing, obtaining, storing, transmitting, disclosing and profiting from biometric identifiers and biometric information. Texas prohibits capturing a biometric identifier for a commercial purpose without informed consent. Washington's regime turns on enrolling a biometric identifier in a database for a commercial purpose. None of these are drafted around the moment a product recognises a customer; they are drafted around the handling of the identifier itself.

A training corpus of face embeddings is handled biometric data. So is a corpus of raw images destined for template generation, once the generation occurs. The reason this feels novel is that the teams building corpora and the teams reading privacy statutes have historically been different people working from different documents — the research side treats a dataset as a technical input and the legal side reviews the deployment.

Where a Product Team Actually Picks This Up

HIGH RISK
Licensing a face-recognition or face-matching model from a vendor that will not describe its training corpora, or that describes them as 'publicly available images'.
HIGH RISK
Fine-tuning an identity model on images collected from your own users for one purpose and repurposed for another, without consent covering the training use.
HIGH RISK
Buying a dataset from a broker where the chain of consent is asserted in a sentence and cannot be traced to a collection event.
MODERATE
Using a research-released academic face dataset in a commercial product. Several widely cited sets carry research-only terms, and some have been withdrawn by their publishers while copies circulate.
MODERATE
Voice models trained on scraped speech. Voiceprints are biometric identifiers in the same statutes, and speech corpora receive a fraction of the provenance scrutiny image corpora now get.
LOWER
Vision models that detect faces without generating identifying templates, and synthetic face datasets generated without conditioning on identified real individuals — though the second requires the generator's own provenance to hold up.

Synthetic is not automatically clean. A synthetic face dataset inherits the provenance of the generator that produced it. If that generator was trained on scraped identified faces, the synthetic set is downstream of the same collection, and a defence resting on the images being fictional has to survive the question of what the model learned identities from. Ask the same provenance questions one layer up; synthetic data moves the question, it does not answer it.

The Remedy That Makes This a Continuity Problem

Ordinary data-privacy exposure is financial: damages, penalties, remediation cost. What distinguishes unlawfully collected training data is that regulators have sought — and settlements have included — destruction of the derived models, not only the underlying records. The logic is straightforward: leaving the model intact would let the party keep the benefit of the collection it was not entitled to make.

For a company that licensed the model rather than built it, this reframes the diligence question. You are not primarily asking whether your vendor can pay a judgment. You are asking whether the capability your product depends on can be ordered out of existence, and what your agreement says happens to you if it is. Very few AI vendor agreements contain a meaningful answer, because the template they were drafted from contemplated a software outage rather than the permanent unavailability of a trained artefact.

The Diligence Set, in Questions You Can Send Today

  1. Name every training corpus. For each: source, collection method, approximate date range, and whether it contains identified individuals. "Publicly available web data" is not an answer to this question.
  2. State the lawful basis per corpus. Consent obtained at collection, a licence from a party who obtained consent, or a claim that the data is not biometric. Each of the three is a different argument and should be identified as such.
  3. Ask about withdrawals. Whether any component dataset has been retracted, restricted to research use, or made subject to a deletion order — and what the vendor did about it.
  4. Ask what deletion means. If a subject exercises a deletion right, what happens to the template, the derived embeddings and the model. A vendor with no answer beyond deleting a row has not thought about the artefact that matters.
  5. Seek a warranty and an indemnity. Lawful collection warranted, biometric and privacy claims indemnified, and no cap so low that the indemnity is decorative.
  6. Plan for withdrawal. An exit or substitution path if the model becomes unavailable. This is a product decision, not a legal one, and it is the one most worth having before you need it.

Your Own Deployment Duties Do Not Go Away

Upstream provenance is the part teams miss; downstream collection is the part they owe regardless. If your product generates a template from a face, voice, iris or behavioural signal belonging to a real person, you are collecting a biometric identifier and the full apparatus applies — written notice before collection, informed consent, a published retention and destruction schedule, and limits on disclosure and on profiting from the identifier. Buying a well-documented model satisfies none of that. The two workstreams are independent, and each is regularly used as an excuse for not doing the other.

Frequently Asked Questions

Our model was trained abroad on data from non-US subjects. Does a US biometric statute still reach us?

The training may sit outside a given statute's reach if no covered individuals' data was involved, but that is a fact question you need evidence for rather than an assumption — large scraped corpora are rarely geographically clean, and a vendor that cannot describe collection cannot support the claim that no residents of a particular state were included. Separately, your deployment reaches whoever your users are: a model trained entirely abroad and used to scan customers in Illinois or Texas triggers the local duties in full.

Does a photograph become biometric data the moment it lands in our storage?

Generally no. Most biometric statutes expressly exclude photographs from the definition of a biometric identifier while covering the scan of face geometry derived from one, and the practical line is whether your system generates an identifying template. This distinction is where features cross without a ticket: a photo library is out of scope until someone enables face grouping, at which point the same bucket of images is producing identifiers and the notice, consent and retention duties attach retroactively to nothing and prospectively to everything.

We use a third-party API for face matching and never store templates ourselves. Are we in scope?

Yes, in most framings. Determining that faces will be matched and directing the processing is what makes you the responsible party; the vendor performing the computation is acting on your instruction. Several statutes also reach disclosure of biometric identifiers to a third party, which is precisely what sending an image to an external API for template generation is — and disclosure typically requires its own consent. The contractual controls that matter are that the vendor processes only on your instruction, does not use your data to train, and deletes on instruction with written confirmation.

How is this different from the copyright fight over training data?

Different right-holder and a different remedy. Copyright claims are brought by rights owners over expressive works and turn substantially on fair use. Biometric claims are brought by or on behalf of the individuals depicted, do not depend on anyone owning the image, and in Illinois carry a private right of action with statutory damages per violation — which is what makes the class arithmetic significant. A licence from a photographer resolves the copyright question and does nothing at all for the biometric one, because the person in the frame was never the licensor.

What single artefact best demonstrates we did the work?

A dataset register: every corpus used to train or fine-tune any model that touches faces, voices or other biometric signals, with source, collection basis, date acquired, licence terms, and a named owner. It is the biometric analogue of a software bill of materials, it answers most of what a customer's security review asks, and it is the only version of this record that survives the departure of the people who built the model. Start it before you need it — reconstructing corpus provenance after the fact is the part that cannot be done.

Build the Dataset Register First

You cannot inspect a model for provenance. Everything you will ever be able to say about where your biometric capability came from has to be written down at acquisition time by the person who acquired it.

List the corpora, record the collection basis, get the warranty and indemnity in the vendor agreement, and keep your own deployment consent and retention duties on a separate track. Those two tracks, run independently, are the whole of a defensible biometric posture.