RatedWithAI

RatedWithAI

Accessibility scanner

AI Legal & ComplianceAugust 10, 2026

Fair Use Protected the Training. It Doesn't Protect Your Output.

Every favorable AI training ruling gets read in marketing meetings as permission to publish. It isn't. The defense belongs to the party that did the training, and the risk in a published output belongs to whoever published it. Understanding where the two questions separate is most of the practical work.

Two questions
Was the training lawful? Does this output infringe? Different defendants.
Factors 1 & 4
Transformativeness and market effect carry the analysis in AI cases
Provenance
How training copies were acquired is a separate exposure from how they were used

The Conflation That Creates the Exposure

Fair use is a defense to copyright infringement. Defenses are personal to a defendant and specific to an act. When a court considers whether ingesting a corpus of books to train a model is fair use, the act under examination is the model developer's copying. The defendant is the model developer. The holding does not transfer to you, and it does not describe your conduct.

Your act is different: you generated something and put it on a website, in an ad, in a shipped product. The question for that act is the ordinary one — does the thing you published copy protectable expression from a copyrighted work? That analysis starts at substantial similarity, and it does not care how the model was trained.

QUESTION 1 — NOT YOURS
Was training on these works lawful?
  • Defendant: the model developer
  • Turns on transformativeness and market harm
  • Separately, on how the copies were obtained
  • Resolved in litigation you are not a party to
QUESTION 2 — YOURS
Does this specific output infringe?
  • Defendant: you, the publisher
  • Turns on substantial similarity to a specific work
  • Unaffected by the training ruling
  • Resolved by a demand letter addressed to you

What the Four Factors Are Actually Testing

The statutory factors are familiar; how they behave in AI disputes is less so. In practice, two of the four are close to dispositive.

Factor 1 — Purpose and characterHIGH
Is the use transformative — does it serve a different purpose than the original, rather than substituting for it? Arguments that a model learns statistical relationships rather than storing expression live here. Commerciality matters but rarely decides.
Factor 2 — Nature of the workLOW
Creative works get more protection than factual ones. In mass-ingestion cases this factor is usually acknowledged and then set aside; the corpus contains everything.
Factor 3 — Amount and substantialityMEDIUM
Training typically involves the whole work, which sounds fatal and is not — copying everything can be reasonable where it is necessary to the transformative purpose. This factor bites hardest on the output side, where reproducing a small but recognizable core is enough.
Factor 4 — Market effectHIGH
Does the use harm the market for the original, including licensing markets? The contested frontier is whether flooding a market with machine-generated substitutes counts as cognizable harm even when no individual output copies anything.

The pattern that has emerged across the first wave of decisions is not "AI training is fair use" or "AI training is infringement." It is narrower and more useful: the more the system substitutes for the original in its own market, and the less lawful the acquisition of the copies, the worse it goes. A tool that summarizes what a document says is on very different ground from one that outputs a competing version of the document.

Acquisition Is Its Own Exposure

One of the most commercially relevant distinctions to come out of this litigation is between using a work to train and obtaining the copy in the first place. A holding that the training use was transformative does not retroactively legitimize a library assembled from pirated files; the downloading can be its own infringement with its own damages. For a buyer, this converts an abstract legal debate into a concrete diligence item: where did the training corpus come from, and will the vendor say so in writing?

Pre-Publication Checks That Actually Reduce Risk

None of this argues against using generative tools. It argues for a small number of checks placed where the risk concentrates — on outputs that are public, commercial, and close to someone's protected expression.

Never prompt with a living creator or brand by name
Style prompts naming a specific artist, studio, photographer, or brand are the single highest-signal fact in a plaintiff's complaint. They also void several vendor indemnities outright. Describe the aesthetic instead of naming its owner.
Reverse-image and similarity search anything customer-facing
For images and logos, run a reverse image search before publication. For code, run license scanning on generated snippets. For long-form copy, run a plagiarism check. These take minutes and catch the memorization cases that matter.
Do not generate anything that will function as a trademark
Logos and brand marks carry trademark risk on top of copyright risk, and you cannot register the machine-generated portions as your own copyright. Commission the marks you intend to enforce.
Keep the prompt and model version with the asset
Provenance records are what let you demonstrate independent creation, and they are what let you invoke a vendor indemnity. Store prompt, model, version, and date alongside the file rather than in a chat history someone will clear.
Read the indemnity's exclusions, not its headline
Assume the indemnity does not apply if you turned off filters, supplied third-party inputs, modified the output, or ran an older model version — then verify. A cap set at fees paid is not coverage against statutory damages across a content library.
Separate the two questions in your own risk register
Training-data litigation is a vendor risk you manage through contracts and vendor selection. Output similarity is a product risk you manage through review. Teams that log them as one line item under-invest in the second one, which is the one addressed to them.

The Asymmetry Worth Planning Around

Machine-generated expression is not protectable as your copyright, but it can still infringe someone else's. You carry the downside without the upside. That asymmetry is a reason to be deliberate about which assets you generate: high-volume, low-stakes content where copying by competitors is irrelevant is a good fit; brand marks, flagship creative, and anything you would want to enforce against a copycat is not.

Frequently Asked Questions

A court held that training on copyrighted books was fair use. Doesn't that settle it?

It settles one question for one defendant on one record. It does not address whether an output you publish is substantially similar to a protected work, and it does not bind courts in other circuits or on different facts. Treat favorable training rulings as reducing vendor risk, not as clearing your publication risk.

Which fair use factors matter most in AI cases?

Factor one, on whether the use is transformative, and factor four, on market effect. Factor three behaves unusually because training typically uses whole works, which courts have accepted where the copying serves the transformative purpose. Factor two rarely moves the outcome in mass-ingestion cases.

Does it matter that our use is commercial?

It is relevant but not decisive. Commerciality weighs against fair use under factor one, yet plenty of commercial uses are fair. What tends to matter more is whether your use substitutes for the original in its own market — a commercial use that serves a different purpose fares better than a non-commercial one that displaces the source.

How was the training data acquired, and why should we care?

Because unlawful acquisition is a separate wrong from the training use, and a favorable ruling on the latter does not cure the former. Ask vendors to describe corpus provenance and to represent that copies were lawfully obtained. Refusal to put it in writing is itself informative.

Can we rely on our vendor's copyright indemnity?

Only within its exclusions and its cap. Common carve-outs cover disabled filters, user-supplied inputs, prompts naming third-party content or styles, modified outputs, and outdated model versions. Map your actual workflow against the exclusion list; teams often discover their standard process falls outside coverage.

Can we copyright the content our AI produces?

Only the human-authored contribution — selection, arrangement, and substantive modification. Purely machine-generated elements are not registrable. Assets you intend to enforce against copying should have a documented human authorship story, or should be commissioned rather than generated.

Put the Review Where the Risk Is

The training-data question will be resolved over years, in courts, by parties that are not you. The output question is resolved every time someone hits publish. A similarity check on public-facing assets and a rule against naming creators in prompts cost almost nothing and remove most of the realistic exposure.

Spend the legal budget on vendor terms and provenance representations. Spend the process budget on the last step before publication.

Related Reading