Can Your AI Vendor Train on Your Data? The Contract Rights Checklist for 2026
Most teams check one sentence — "we don't train on your data" — and stop. The permission that actually governs is usually four clauses apart from that sentence, in a document with a different name, and it is rarely about training at all. It is about what counts as your data once the model has finished with it.
The Risk Is Not Where People Look For It
The frontier model providers are, by 2026, the easy case. Their business and API tiers have converged on not training on customer inputs by default, and they say so in a findable place. The exposure has moved one layer up the stack:
- The AI feature in a tool you already bought. Your support desk, CRM, ATS or meeting recorder shipped an AI feature under the existing MSA. The licence you granted years ago to "host, copy, process and display Customer Data to provide and improve the Services" is now doing work nobody negotiated it to do.
- The vertical AI startup. Small vendors need a data moat and their terms say so. This is not bad faith — a two-year-old company's only defensible asset may be tuned weights built from early customers.
- The free tier your team actually uses. Procurement approved the enterprise plan; three people are pasting into the consumer product because it is faster. The contract you negotiated does not govern that traffic.
Clause One: The Training Permission
Find the licence grant over Customer Data and read the purpose limitation at the end of it. The difference between two words is the whole deal:
Clause Two: Derived Data — Where the Restriction Leaks
A vendor can honestly promise never to train on Customer Data and still train on everything made from it. The mechanism is a definition, usually parked in the definitions section rather than the AI section:
"Derived Data" means data, statistics, metrics, and insights generated by Provider from Customer Data in aggregated or de-identified form. Provider owns all Derived Data and may use it for any purpose, including improving its products, without restriction, during and after the Term.
Applied to an AI product, that definition plausibly covers embeddings of your documents, entities extracted from your tickets, labels your team corrected, prompt and response pairs stripped of identifiers, and evaluation sets built from your edge cases. Those artefacts are frequently more valuable for training than the raw records.
The fix is small and usually accepted: exclude from Derived Data anything that incorporates the substance of Customer Data, and cap the grant at aggregated statistics that do not permit reconstruction of individual records or customer-specific content. Vendors who genuinely only want usage metrics do not fight this.
Clause Three: Who Owns the Tuned Model
Standard allocation looks fair on its face and is not symmetric in effect:
- Customer owns Customer Data. Uncontroversial, and worth less than it sounds.
- Customer owns, or is assigned, Output. Common in 2026 terms; note it is often subject to a caveat that other customers may receive similar output.
- Provider owns the Services, the models, and all improvements, modifications and derivative works. Fine-tuned weights are improvements.
So the customer who supplies the differentiating data and pays for the tuning run ends up with a licence to use a model the vendor owns and can, absent an exclusivity term, offer to that customer's competitors. If the tuned model is a genuine competitive asset, there are three asks in increasing order of difficulty:
- Exclusivity. The tuned artefact is served only to you, and its weights are not merged into models served to others. Cheapest to get.
- Portability. On exit, you receive the tuned weights or an equivalent artefact in a usable format. Feasible with open-weight bases, usually impossible on a closed hosted base — ask early, because the answer determines the base model choice.
- Assignment. You own the tuned artefact outright. Rare, priced accordingly, and mostly relevant where you funded the tuning as a development project rather than a subscription feature.
Clause Four: Exit, and What Deletion Can Honestly Mean
Deleting rows is routine. Removing the influence of those rows from trained weights is not, and a vendor that promises full unlearning on 30 days' notice is either using a narrow technical method it should be willing to describe, or is promising something it will not do. Write exit terms around what can be verified:
- Deletion of Customer Data and all copies, including backups, on a stated schedule.
- Deletion of embeddings, indexes, caches and evaluation sets built from it.
- Deletion of fine-tuned checkpoints created for or from your data.
- A forward-looking prohibition: no further use of any model artefact derived from Customer Data for any customer other than you.
- Written certification of deletion, signed, within a fixed number of days.
- Survival: the restriction outlives the term. Most training permissions are drafted to survive; the corresponding restriction usually is not, unless you say so.
Subprocessors are the hop everyone skips. Your vendor's zero-training promise binds your vendor. The model provider underneath, the vector database, the evaluation tooling and the observability platform each hold your content under their own terms. Ask for the subprocessor list, then ask which of them receive prompt and response content rather than metadata. The answer is frequently more than one.
The Six Questions to Send Before Signing
- Does any agreement, including the base MSA, permit use of our data or anything derived from it to train, tune or evaluate models serving other customers? Cite the clause.
- How is Derived Data defined, and does it include embeddings, extracted entities, or corrected labels from our content?
- Who owns weights produced by tuning on our data, and will you commit that they are served only to us?
- Which subprocessors receive prompt and response content, and what are their retention and training terms?
- On termination, what exactly is deleted, on what schedule, and will you certify it in writing?
- Which of these commitments survive termination, and which are settings you can change unilaterally?
Question six is the one that separates a contract term from a marketing position. A vendor that will put its training restriction in the agreement rather than the trust centre is telling you something real; a vendor that will not is also telling you something real.
Frequently Asked Questions
Our vendor's trust page says they never train on customer data. Isn't that enough?
A trust page is a statement of current practice, not a contractual commitment, and it can be edited without notice or breach. Ask for the same sentence in the agreement or an addendum, with a survival clause. If the vendor is unwilling to contract to what it publishes, treat the published position as accurate today and unenforceable tomorrow — which is exactly the risk profile you are trying to remove.
We're a small buyer with no negotiating leverage. What's worth asking for anyway?
Two things, both cheap for the vendor to grant. First, tighten the Derived Data definition to exclude anything incorporating the substance of your content — most vendors only wanted usage metrics and will agree. Second, make the training restriction survive termination. Ownership of tuned weights and portability are the expensive asks and are usually not winnable below a certain contract value; the definitional fixes are winnable at any size and remove most of the exposure.
Does it matter if our data is de-identified before training?
It reduces privacy exposure and does very little for competitive exposure. A de-identified corpus of your support tickets still teaches a model your product's failure modes, your customers' objections and your team's resolutions — that is the asset, and stripping names does not remove it. Separate the two risks when you negotiate: privacy terms handle personal data, and a confidentiality or competitive-use restriction handles the business content. De-identification is only an answer to the first.
What if we're the vendor and our terms currently permit training?
Expect to be asked about it, and decide deliberately rather than by inheritance. Many AI startups carry an aggressive training grant copied from an early template that no longer matches how they operate, which converts every enterprise security review into a negotiation and costs deals. If you do need training rights, scope them narrowly — opt-in, specified categories, no competitor exposure — and say so plainly. A narrow grant you can defend closes faster than a broad one you have to explain.
How does this interact with our own customer commitments?
It is where the real liability lives. If your terms promise your customers that their data is used only to serve them, and your AI vendor's terms permit training on what you pass through, you have made a promise you are not in a position to keep. Map the chain before adding an AI feature: what you promised upstream must be at least as restrictive as what your vendor grants itself downstream. This is the most common way a routine feature launch turns into a breach of an existing contract.
Read the Definitions, Not the AI Section
Every AI vendor contract has a section that sounds like it answers this question, and the answer is almost never there. It is in the licence grant's purpose limitation, in the definition of Derived Data, and in the ownership clause's word "improvements." Three sentences, none of which mention artificial intelligence.
Pull those three out of your next AI contract before anything else. If they line up, the rest of the review is ordinary procurement. If they do not, no amount of security questionnaire will fix it.