You Called It an A/B Test. The Common Rule Calls It Research.
The experiment ships on a Tuesday. Six months later someone wants to write it up, a reviewer asks for the ethics approval number, and a question nobody asked in advance becomes unanswerable in retrospect.
AI teams publish. The field's hiring, credibility and recruiting run on papers, preprints and blog posts making claims about how people behave with these systems — and the moment an internal experiment is reframed as a contribution to what is known generally, it satisfies the first half of a definition that most product teams have never read. The second half is satisfied almost automatically, because the underlying data is conversation.
Both Gates, or Neither Applies
A systematic investigation, including development, testing and evaluation, designed to develop or contribute to generalisable knowledge.
- +Intent to publish, present or post a preprint
- +Hypothesis stated in advance and tested against a control
- +Conclusions framed as being about people in general, not about your users
- +A named collaborator whose output is an academic paper
- −Optimising a flow to raise your own conversion rate
- −Internal quality evaluation with no external claim
- −Operational monitoring and incident analysis
A living individual about whom the investigator obtains information through intervention or interaction, or obtains identifiable private information.
- +Manipulating what a user sees in order to measure their response
- +Surveys, interviews and moderated sessions
- +Analysing conversation transcripts tied to accounts
- +Re-identifiable behavioural traces, however indirect
- −Benchmarks run against models with no human data
- −Genuinely de-identified aggregate statistics
- −Public documents about organisations rather than people
The useful consequence of a conjunctive test is that one honest answer can resolve it. An experiment whose findings genuinely stay inside the company clears gate one; a study that truly touches no identifiable human data clears gate two. Most disputes are about teams asserting both while planning to publish.
What the Team Says, What the Definition Asks
Is anyone going to write about what it showed? Volume is irrelevant to the definition; the intended contribution to generalisable knowledge is the whole of gate one.
Could identity readily be ascertained from what remains? Free-text prompts routinely carry employer, location, health and family detail that no ID-stripping touches.
Were the required elements of informed consent present, and was refusal genuinely without penalty? Acceptance of terms as a condition of access answers neither question.
Is your company engaged in the research — obtaining data, interacting with participants, or receiving identifiable information? If so you have your own obligations, not a borrowed exemption.
Where does the knowledge land? The distinction is defensible when findings stay internal and operational, and collapses the moment the same analysis becomes a public claim about human behaviour.
Who determined that? Minimal risk changes the level of review and may support an exemption, but that determination is made by a review body, not by the investigator who wants the answer.
Engagement Is How the Obligation Reaches You
Companies routinely assume a university collaborator absorbs the compliance work. Sometimes that is right and sometimes it is exactly backwards. The concept that controls the answer is engagement: an entity is generally engaged in research when its own employees or agents intervene or interact with participants for research purposes, or obtain identifiable private information for those purposes. A company that runs the experiment inside its own product, holds the raw transcripts and hands a partner an analysis is doing far more of the regulated activity than the partner is.
The practical version of this is a written determination made before data collection — who is engaged, whose review covers what, and whether a reliance arrangement is needed — rather than an assumption resolved by whoever wrote the acknowledgements section.
Six Designs That Deserve Review Regardless
Even where the formal regulation does not bind you, a serious deployment review is worth running on the categories below — because the reputational and litigation risk does not track the funding source, and because the first external question after an incident is invariably who reviewed this before it ran.
The Cheapest Version of Getting This Right
A one-page determination form attached to any experiment where publication is plausible, answering both gates and naming who decided, costs a team almost nothing and resolves the entire question at the only moment it is cheap to resolve. Independent review boards will review commercial protocols for a fee where a formal record is needed. Both options are dramatically less expensive than the alternative, which is discovering at submission that a year of work is unpublishable and that the dataset behind it was assembled under a consent theory nobody would defend in writing.
Related Reading
- FDA device status for AI clinical decision support — the regime that attaches when a health-adjacent study becomes a health claim.
- Training-data collection and legal risk — the acquisition-side question about the same corpora.
- Privilege and work-product waiver from AI tool use — another place where an internal artefact becomes externally visible later.
Check What Your Site Says About User Data and Research
"We never use your conversations", "opt out anytime" and research-programme pages accumulate across docs, trust centres and old launch posts — and they are the record against which a consent theory gets judged.
See every claim your site makes in one pass. Run a free scan and check each against what your experiments actually do.
This article is general information and not legal or regulatory advice. The scope of federal human-subjects regulation depends on funding, institutional assurances and the specific facts of a study, and institutional policies frequently impose more than the regulation requires. Consult qualified counsel or a review board before relying on any conclusion here.