Your AI Tuned the Sanctions Screen Down. That Was a Legal Decision, Not an Ops One.
Sanctions liability does not require intent, knowledge or negligence. So when a model decides which of nine thousand name matches deserve a human's attention, it is not managing a queue — it is setting the boundary of a strict liability exposure, and the record of how it was set is the only defense anyone will read.
The Regime Is Unusual, and That Is the Whole Point
Most compliance exposure a software company carries is fault-based in some way. You are liable because you were unreasonable, because you failed to disclose, because you knew or should have known. Sanctions is not built that way. Civil enforcement of US sanctions programs administered by the Office of Foreign Assets Control operates on a strict liability basis: a US person who deals with a blocked party has violated the prohibition, whether or not anyone at the company had any idea.
This changes what a compliance program is for. It is not there to prevent liability, because liability can attach anyway. It is there to prevent the violation from happening, and — when one happens regardless — to change the penalty. OFAC's published enforcement guidelines set out the general factors it considers in determining administrative action, and the existence, nature and effectiveness of a compliance program sits among them, alongside whether the conduct was voluntarily self-disclosed. The program is the mitigation.
Which means the artifact that matters is not the screening tool. It is the evidence of how the screening tool was designed, tested, tuned and supervised.
What AI Actually Changed About Screening
Name screening has always been fuzzy matching. Sanctioned parties appear under transliterations, aliases, name orders that vary by culture, and deliberate obfuscation. A pure exact-match screen catches almost nothing useful, so screening engines have long produced a similarity score and a threshold above which a human reviews the hit.
The change is that the scoring function is increasingly a learned model rather than an edit-distance formula, and that a second model is often placed on top of it to triage, auto-close or summarize alerts before a human sees them. Both are defensible choices. Both create a specific problem: the reason a particular alert was closed is now much harder to state in a sentence.
- Auto-closure is the sharp edge. An alert surfaced and dismissed by a human leaves a rationale. An alert never surfaced leaves nothing at all. If a model closes alerts without human review, the closure logic itself has to be documented, tested and periodically sampled by a person.
- Score drift is silent. A learned matcher retrained on recent data can quietly shift what a score of 0.85 means. A threshold that was validated against the old distribution is no longer the threshold anyone approved.
- Explanations are not evidence. A generated summary of why an alert looked benign is a description, not a record of the underlying comparison. Keep the inputs, the list version, the score and the decision, not only the prose.
- Vendor opacity becomes your gap. If you cannot describe, at least at a functional level, what your screening provider's model keys on and how it is tuned, you cannot demonstrate that you tested it.
Threshold Tuning Is the Document Everyone Will Ask For
Every screening program eventually faces the same pressure. The alert queue is unmanageable, the overwhelming majority of hits are false positives, and someone proposes raising the match threshold. This is a normal and often correct operational decision. It is also, viewed from the outside afterward, the moment the company decided how many true matches it was prepared to miss.
The difference between a defensible tuning change and an indefensible one is entirely in the paperwork surrounding it.
- Pre-change testing against known positives. Run a set of true matches — including hard ones with transliteration and alias variants — through both the current and proposed configuration. If the proposed threshold drops any of them, that is the finding.
- A stated rationale that is not only volume. "The queue was too large" is a reason to add reviewers or improve data quality. It becomes a reason to change detection only when paired with evidence that detection is preserved.
- Named approval above the operating team. The person who owns the alert backlog should not be the sole approver of the change that shrinks it.
- Post-change validation on a schedule. Re-test after the change is live, and again periodically. A one-time validation supports a one-time claim.
- Version everything. Which model version, which list version, which threshold, effective from when. Reconstructing the configuration in force on a date eighteen months ago is a routine request and an unroutine amount of work if nobody planned for it.
The Non-Financial Company Problem
Sanctions screening is culturally associated with banks, and that association causes a specific failure. Software companies, marketplaces, agencies and infrastructure providers frequently have no screening at all, on the theory that sanctions are a financial services concern.
Sanctions prohibitions reach US persons broadly and are not limited to moving money. Providing services to a blocked party, exporting software, paying an overseas contractor, or onboarding a customer whose ownership traces back to a designated entity can all raise the question. The ownership point catches people out most often, because a counterparty that is not itself listed can still be blocked by operation of ownership rules when designated parties hold a sufficient interest — which is not visible from the company name.
The correct response is not necessarily to build a bank-grade program. It is to do a documented risk assessment: where are our customers, who do we pay, what do we deliver, what is the realistic exposure, and what controls are proportionate. A reasoned conclusion that formal screening is not warranted is a defensible position. An unexamined assumption is not the same thing, and it is what an investigation finds.
A Screening Program a Reviewer Can Follow
- Written management commitment and a named owner. OFAC's compliance framework treats senior management commitment as foundational, and it is the element most commonly absent at smaller companies.
- A risk assessment that drives the controls. The screening scope should follow from the assessment, and the assessment should be revisited when the business changes — new geographies, new payment flows, new reseller channels.
- Screening at onboarding and on list updates. Designations take effect on publication. A customer screened clean in March is not screened for a May designation unless you rescreen the base.
- Independent testing of the automated component. Someone who does not own the tool should periodically test it, including sampling auto-closed alerts.
- Training that reaches sales and support. The people most likely to encounter the first signal — an odd payment routing, a request to ship somewhere unexpected, a customer asking to change the invoicing entity — are usually not in compliance.
- An escalation path that assumes urgency. If a potential hit is real, the transaction needs to stop and counsel needs to be involved that day, not at the next weekly review.
Frequently Asked Questions
If our screening vendor's model missed a designated party, can we point at them?
Not to the regulator. The prohibition binds the party engaging in the transaction, and using a service provider does not transfer that. A vendor agreement may give you a contractual remedy for cost, and it is worth negotiating one, but the enforcement conversation is with you. This is also why diligence on the provider — what it screens, how often lists refresh, what it auto-closes — is itself part of your program.
How do we handle the sheer volume of false positives without weakening detection?
Attack data quality and context before attacking the threshold. Most false positive volume comes from screening thin records: a name with no date of birth, no country, no identifier. Adding secondary attributes to the comparison usually cuts noise far more than raising a score cutoff, and it does so without changing what the system is capable of catching.
Is it acceptable for a model to close alerts with no human review at all?
It is a risk decision rather than a categorical prohibition, and it depends heavily on what is being auto-closed and on what evidence. If you do it, the safeguards that make it defensible are a documented and tested closure rule, a conservative scope, human sampling of closed alerts on a defined cadence, and retention of the underlying comparison data so a closed alert can be reconstructed later.
We are a small SaaS company with international customers. What is the minimum?
A documented risk assessment, screening of customers and vendors against the relevant lists at onboarding with periodic rescreening, a defined escalation path to counsel, and retention of what was screened and when. That is a materially smaller lift than a bank program and it addresses the failure mode where nobody ever looked at the question at all.
What should we retain, and for how long?
Retain the screening inputs, the list version and date, the match scores, the disposition and reviewer, and the configuration in force. Sanctions recordkeeping requirements are longer than many teams assume and the practical constraint is worse than the legal one, because reconstructing a historical configuration from a system that overwrites it is not possible. Confirm the applicable retention period with counsel for your specific programs.
Does an EU or UK sanctions exposure work the same way?
The general shape is similar — list-based prohibitions, ownership and control rules, expectations of a proportionate program — but the specifics of liability standards, licensing and reporting differ meaningfully by regime, and a company with cross-border operations is usually subject to several at once. Screening against a single list because it is the one your vendor defaults to is a common and avoidable gap.
What Does Your Site Promise About Compliance?
Screening and compliance claims tend to accumulate across a marketing site — a trust page, a security page, a comparison table, an old launch post — written at different times by different people, and rarely read together against what the product actually does today.
See every compliance and capability claim on your site in one pass. Run a free scan and check them against your current controls.
This article is general information about a complex regulatory area and is not legal advice. Sanctions programs differ by jurisdiction and change frequently. Consult qualified counsel about your specific facts, and involve counsel immediately if you believe a potential violation has occurred.