You Can Delete the Data. You Cannot Delete It From the Weights.
Standard technology diligence confirms the target owns its code and holds its customers. It does not confirm the model may lawfully remain trained on what it was trained on — and that is the finding that arrives after closing, when it is the buyer's problem.
Why the Usual Workstreams Miss It
Technology diligence is organised around artefacts that have owners: repositories, dependencies, patents, trademarks, and customer contracts. A trained model is a different kind of asset. Its value derives from data the target may or may not have had the right to use, under permissions granted by parties who may since have withdrawn them, in a form from which individual contributions cannot be extracted.
None of the standard workstreams owns that question. IP diligence looks at what the target created; privacy diligence looks at how data is stored and disclosed; security diligence looks at controls. The question of whether the model can lawfully stay as it is falls between them, which is why it survives to post-closing so often.
The Four Findings That Change Price
The AI Diligence Request List
- ☐An inventory of every dataset used in training or fine-tuning, with source and licence for each
- ☐The contractual basis for using customer data in training, quoted from the actual agreements in force at the time
- ☐A log of deletion, correction, and opt-out requests, and how each was reflected in deployed models
- ☐Any dataset acquired from a third party, with the terms permitting onward use by an acquirer
- ☐Every base or foundation model in the stack, with the exact licence text and version
- ☐Acceptable-use restrictions, user thresholds, and derivative-model clauses in each licence
- ☐Whether the licence permits transfer on a change of control without consent
- ☐Upstream API provider terms, including data-use, indemnity, and termination rights
- ☐A list of every automated decision affecting an individual, whether or not it is branded as AI
- ☐Assessments, audits, or disclosures already performed, with dates and who performed them
- ☐Regulatory correspondence, complaints, or claims touching an automated system
- ☐Commitments made to customers about AI behaviour in contracts, security questionnaires, or marketing
- ☐Specific representations tied to each finding, not generic AI-compliance boilerplate
- ☐A special indemnity where a provenance defect is identified but not resolved pre-closing
- ☐Closing conditions requiring documented remediation of unsatisfied deletion obligations
- ☐Price adjustment where retraining is the only available remedy, scoped to its actual cost
What a Good Target Looks Like
The distinguishing feature is not the absence of problems; it is the presence of records. A target that can produce a dataset registry, a licence file per source, and a mapping from customer agreements to training runs has usually thought about the questions above before you asked them, and its answers can be verified rather than taken on faith.
A target that cannot produce those records is not necessarily non-compliant, but the buyer has no way to establish that it is compliant either — and the representation the seller offers is then the only protection, backed by whatever escrow the deal carries. That is a materially different risk profile, and it belongs in the price rather than in a footnote.
Frequently Asked Questions
How early should the AI workstream start?
At the same time as IP diligence, because the answers feed the same conclusions about what the buyer is actually acquiring. Starting it late is the common pattern and the expensive one: provenance findings can require retraining cost estimates, and those take time to produce, so a late start either delays signing or pushes an unpriced risk into the representations.
Does a target's ISO 42001 certificate or NIST AI RMF mapping shorten diligence?
It shortens the documentation hunt considerably, because both frameworks generate exactly the artefacts a buyer needs to review. Neither is a substitute for reviewing the underlying facts — a management-system certificate attests to process, not to the lawfulness of any particular dataset — but a certified target usually hands over an organised file instead of a reconstruction exercise.
What about acqui-hires where the model is not the point?
The diligence is lighter but not absent. Even where the buyer intends to retire the product, retained customer data, outstanding deletion obligations, and any pending claim relating to automated decisions travel with the entity. Confirm what is being assumed and what is being wound down, and make sure the wind-down actually discharges the obligations rather than leaving them dormant.
Is an indemnity enough to cover a provenance defect?
It helps, but it is not equivalent to a fix. An indemnity converts the risk into a claim against escrow or the seller's balance sheet, which is only as good as the amount, the survival period, and the counterparty. A defect whose remedy is a full retraining run can exceed a typical escrow, and it also costs the buyer time and model quality that no payment restores. Where the defect is identified pre-closing, remediation as a closing condition is the stronger structure.
Related Guides
Ask for the Mapping, Not the Assurance
Every question in this list has a documentary answer or it does not. A target that can map customer agreements to training runs has an auditable position; one that offers a verbal assurance has an untested one, and the difference belongs in the price.
Buyers of consumer-facing products should pair this workstream with an accessibility review of the target's live interfaces, since inherited exposure there is straightforward to measure before signing and expensive to discover afterwards.