The Sandbox Is a Supervised Runway, Not an Exemption
Two things get confused constantly. A regulatory sandbox is a place to build a high-risk system under a regulator's eye before it goes to market. Testing in real-world conditions is a separate permission to put an unfinished system in front of actual people. Neither suspends a single requirement, and the second one carries obligations most teams discover after the pilot has already started.
What a Sandbox Is For
Member states are required to establish at least one AI regulatory sandbox at national level, alone or jointly with others, and to have it operational. The design goal is narrow and worth stating plainly: give providers a controlled environment to develop, train, validate and test an AI system before it is placed on the market, with the competent authority supervising and advising as it goes.
That is the whole offer. There is no derogation attached. If your system is high-risk, it still needs its risk management system, its data governance, its technical documentation, its logging, its human oversight design and its conformity assessment before it can be sold. The sandbox changes when the regulator sees that work, not whether you have to do it.
The Three Things Admission Is Actually Worth
- Guidance you can rely on. Supervision and support during development means questions that would otherwise sit unanswered — is this use case in Annex III, does this data practice hold up, is our oversight design adequate — get an authority's view while the architecture is still movable.
- An exit report that counts as evidence. On completion, the authority produces a written record of the activities carried out, the results and the learning outcomes. That document supports the conformity work and can be shown to notified bodies, buyers and investors. Most teams generate nothing comparable on their own.
- Fine relief for what you surface. Where you respect the sandbox plan and follow the authority's guidance in good faith, no administrative fines are imposed for infringements identified during participation. That is a meaningful incentive to bring the awkward parts of the system in rather than to keep them out of scope.
None of this touches civil liability. Harm caused during participation is still actionable under ordinary liability law, and a sandbox letter is not a defence to a damages claim.
Testing in Real-World Conditions Is a Different Instrument
The regime that catches people out is testing a high-risk system in real-world conditions outside a sandbox. Teams treat this as an ordinary pilot because it looks like one commercially. It is not, and the conditions are specific:
- A testing plan, submitted. Drawn up and submitted to the market surveillance authority in the member state where the testing happens, with the ability for the authority to approve, request changes, or object.
- Registration before you start. The testing is registered in a dedicated section of the EU database with a single identification number, which is what turns the pilot into a visible, supervisable activity.
- Informed consent from subjects. Freely given, before participation, with information about the nature and objectives of the testing and the conditions of participation, and a right to withdraw without detriment and without justifying the decision.
- Human oversight with real authority. Carried out by people who are suitably qualified and have the means to oversee it — which includes the ability to reverse or disregard the system's outputs.
- A time limit. Six months as the baseline, extendable once by a further six where the extension is justified and notified.
- Serious incident reporting. Incidents during testing are reportable to the market surveillance authority, with immediate suspension where mitigation is not possible.
Read that list against how a typical design-partner pilot is run — a signed order form, a Slack channel, and a promise that the model is "still learning" — and the gap is obvious. The order form is not consent to be a test subject, and the customer's employees whose outcomes the system affects are usually not parties to it at all.
Where the Two Regimes Meet
Real-world testing can also be run inside a sandbox, under the sandbox plan and the supervising authority. The trade is straightforward: more oversight and more paperwork up front in exchange for a defensible position and a written record at the end. If your buyers are regulated — banks, insurers, hospitals, public bodies — that record is worth more than the weeks it costs, because their own assessments depend on documentation that only you hold.
If your system is not high-risk, most of this is optional, and the honest answer is that a sandbox may not be worth the cycle time. The decision hinges on classification, which is the step to get right first.
A Decision Sequence That Avoids the Expensive Mistake
- Classify before you plan anything. High-risk or not, provider or deployer. Every downstream question — sandbox, testing regime, registration — is answered differently depending on the result.
- Ask whether real people are exposed. If your validation runs on historical data with no live effect on anyone, you are not doing real-world testing and the regime does not attach. If a live output changes what happens to a person, it does.
- Check the national scheme, not the summary. Sandboxes are established member state by member state, and the application windows, sector focus and eligibility criteria differ. Pick the one aligned to your first market rather than the first one you find written up in English.
- Budget the calendar, not the fee. For EU SMEs the access is free; the cost is the elapsed time in application, plan agreement and supervision. Run it in parallel with engineering rather than after it.
- Write the consent flow early. Retrofitting informed consent, withdrawal and a data-deletion path onto a live pilot is substantially harder than designing them in, and withdrawal without detriment has product consequences.
- Keep the exit report. It feeds the technical documentation, shortens the buyer questionnaire, and is the single most reusable artefact the exercise produces.
Common Questions
Can we say we are 'EU AI Act approved' after a sandbox?
No, and the claim is the kind that draws attention. A sandbox exit report describes activities and outcomes; it is not a conformity assessment, a CE marking or an authority endorsement of the product. Describe what actually happened — supervised development with a named authority — and let the specificity do the work.
Does the sandbox cover data protection too?
Participation does not suspend the GDPR. There is a narrow provision allowing further processing of lawfully collected personal data for developing certain AI systems in the public interest inside a sandbox, under strict conditions and safeguards. That is a specific carve-out with conditions attached, not a general research licence, and the data protection authority remains involved.
We are a deployer, not a provider. Is any of this relevant?
Real-world testing conditions apply to prospective providers, but deployers frequently host the testing. If you are the organisation whose customers or staff are exposed, you want the plan, the registration number and the consent mechanics in the contract, because the affected people are yours and the reputational exposure lands on you first.
What happens if we run a live pilot without any of this?
You are placing an unassessed high-risk system into use, which is the ordinary infringement — with none of the sandbox fine relief, because that relief attaches to supervised participation. Uninformed subjects and unregistered testing also make the position harder to remediate later, since consent cannot be granted retroactively.
Is a sandbox worth it for a small vendor with one EU customer?
Often not, if your system is not high-risk. If it is, the calculus changes: the exit report and the documented authority engagement tend to be the difference between a procurement cycle you win and one that stalls in a risk committee for two quarters.
Can we withdraw once we are in?
Yes — participation is voluntary and can be ended, and the authority can also suspend or terminate where risks cannot be mitigated. Plan for the possibility that the supervision surfaces something you have to fix before launch. That outcome is the mechanism working, not a failure of the exercise.
What Your Site Claims About Testing Is Discoverable
"Beta", "pilot", "learning from your data" and "approved" all read differently to a regulated buyer than they do to your marketing team — and AI assistants repeat that wording verbatim when they summarise your product to a prospect.
Run a free scan of your site to see what you are currently telling buyers about how your AI is tested.