Your Competitor Can Legally Copy It: What You Actually Own in AI-Generated Content
The entire AI copyright conversation runs one direction — the risk that your output infringes somebody. Turn it around. You published two hundred machine-written pages and a rival lifted forty of them verbatim. What is the claim?
The Asymmetry Nobody Priced In
A company that scales content production with generative tools sits on both sides of an uneven trade. On the input side it carries real exposure: the model was trained on material it did not license, output can reproduce protected expression, and the indemnities offered by vendors are narrower than the marketing suggests. On the output side it acquires substantially less than it thinks — because US copyright protects works of human authorship, and a passage no human composed has no author to hold the right.
The Copyright Office has applied this consistently through its registration decisions and policy guidance: material generated by an AI system in response to a prompt is not registrable, applicants must disclose and disclaim AI-generated content, and what remains registrable is the human contribution. Courts addressing the question have reached the same place on the authorship requirement.
For most businesses this is abstract until the day a competitor republishes their comparison pages. Then the general counsel asks what the claim is, and the answer depends entirely on decisions made months earlier in the content workflow.
What Survives: Four Layers That Still Bite
1. Human-authored contributions
Anything a person actually wrote carries ordinary copyright, and it does not stop being protected because it sits next to machine-generated text. Original analysis, a distinctive argument, commentary drawn from the author's own experience, and human-written framing are all protectable. The practical consequence is that a page built as a human core with AI-assisted expansion has a claim, and a page built as raw generation with light cleanup does not.
2. Selection, coordination, and arrangement
Compilation copyright protects original selection and arrangement even where the underlying elements are unprotected — the settled principle that a factual database earns no protection for the facts but can earn it for a creative structure. A methodology for which items you cover, how you sequence them, and what comparison dimensions you chose can be protectable, and it is often the thing a copyist actually took. It is thin protection and it does not stop paraphrase.
3. Contract
Terms of use bind whoever accepted them, and they can restrict copying, scraping, and redistribution independent of any copyright. This is why the strongest position for high-value material is behind authentication with accepted terms rather than on an open page. Breach of contract against an accepted-terms party is a materially better claim than a copyright claim you may not have, and it does not require proving authorship at all.
4. Trade secret
Material that is never published, derives value from not being generally known, and is subject to reasonable secrecy measures can be protected as a trade secret regardless of who or what produced the text. Internal pricing models, proprietary benchmark data, and unpublished playbooks fit. The moment it goes on a public page, this layer is gone — which is a reason to think carefully about what actually needs to be indexed.
Trademark and passing off reach further than people expect
A competitor who copies a page usually copies more than the prose. Brand names, product names, distinctive visual identity, and proprietary scoring or certification marks travel with the text, and those are trademark questions with no human authorship requirement. Where the copyist presents the material as their own original research, false advertising and unfair competition theories can also apply. In practice, the demand letter that works is frequently not a copyright letter — it is a trademark and unfair competition letter with the copying as context.
The DMCA Problem
A takedown notice is the reflex response to scraped content, and it is the point where the authorship gap becomes concrete. The notice requires a statement, under penalty of perjury, that the sender is authorized to act on behalf of the owner of an exclusive right, and a good faith belief that the use is not authorized by law. If the material is entirely machine-generated, the exclusive right underpinning both statements is doubtful.
Most notices are honored without scrutiny, so this rarely surfaces. When it does — a counter-notice, a sophisticated recipient, a misrepresentation claim — the sender is defending a sworn statement about ownership of content they know a model wrote. That is an avoidable position.
The workable approach is to ground the notice in identified human-authored elements and original arrangement, and to name them specifically rather than asserting blanket ownership of the page. That requires knowing which parts of your own content are human-authored, which is a records problem long before it is a legal one.
Building a Content Operation That Owns Something
1. Design for a human core, not a human polish
Decide before drafting which parts of a page carry the human authorship: the thesis, the analysis, the recommendation, the original data interpretation. Generate around that core rather than generating the whole and editing down. Same output volume, materially different rights position.
2. Keep contemporaneous authorship records
Version history that shows what a person wrote versus what was generated is the evidence that makes a claim provable, and it is nearly impossible to reconstruct later. Editorial tooling that preserves revision history by author does most of this automatically if the workflow does not flatten drafts into a CMS paste.
3. Register the works that matter, and disclaim correctly
Registration is a prerequisite to filing an infringement suit for US works and unlocks statutory damages and fees when timely. Applications must disclose AI-generated material and claim only human contributions; a registration that overclaims is vulnerable, and an inaccurate one can be challenged outright.
4. Write terms of use that stand on their own
Prohibit scraping, bulk copying, and redistribution, and make acceptance meaningful for gated material. Contractual restrictions do not depend on authorship, which makes them the most reliable layer for content produced at machine scale.
5. Decide deliberately what stays unpublished
Publishing converts a potential trade secret into unprotected material in exchange for reach. That trade is often worth it, but it should be a decision. The proprietary dataset behind the article can frequently stay behind the article.
6. Read your vendor terms for what you get, not just what you risk
Provider terms typically assign whatever rights they can in output and disclaim that any rights exist. They cannot manufacture a copyright the statute does not recognize. Read the output-ownership clause as a promise not to compete with you over the material, not as a grant of exclusivity against the world.
Frequently Asked Questions
Our AI vendor's terms say we own the output. Doesn't that settle it?
It settles the relationship between you and the vendor — they will not claim the material and will not license it out from under you. It cannot create an exclusive right against third parties, because ownership clauses transfer rights that exist and copyright in purely machine-generated expression does not. Read those clauses as a non-assertion covenant, which is genuinely useful, rather than as a property grant.
Does this differ outside the United States?
Yes, and it matters for multi-market publishers. The UK provides for computer-generated works with no human author and assigns authorship to the person who made the arrangements necessary for creation, a regime with no US analogue, though its scope is under active debate. Most EU member states apply an own-intellectual-creation standard that functions similarly to the US human authorship requirement. A single portfolio can be protected in one market and unprotected in another.
What about images, video, and audio we generate?
The same authorship analysis applies to the generated elements. Human contribution can still create protection — original composition choices a person made and executed, human-created elements combined with generated ones, and creative arrangement in an edit. Registration practice again expects the generated components to be disclaimed. The asset most companies over-assume rights in is the generated hero image sitting on every page.
Someone copied our AI-written pages and now outranks us. Any recourse?
Search ranking is a platform question rather than a legal one, and duplicate-content handling is the first place to look. On the legal side, work through the layers: human-authored elements, original selection and arrangement, any accepted terms of use, trademark elements copied along with the text, and whether the copyist is representing the work as their own. Frequently the strongest available letter is not the copyright letter.
Is there a threshold of human editing that makes content safe to claim?
No bright line exists, and treating one as though it does is the risk. Protection covers the specific human contributions rather than attaching to the whole document once a quota is met. The useful discipline is being able to point at particular passages and structural choices and say a person authored those — if you cannot do that for a given page, assume the page is copyable.
Volume Is Not an Asset If Anyone Can Take It
Content operations are usually justified as building an asset. An asset is something you can stop other people from using. A library of purely machine-written pages is a traffic acquisition cost that any competitor can copy at zero marginal effort, and the difference between that and a defensible portfolio is decided in the drafting workflow rather than in a legal review afterward.
Put a human core in the pages that matter, keep the records that prove it, register the ones worth registering, and put terms around the rest. None of it slows production much. All of it decides whether you have a claim on the day you need one.