NAIC AI Model Bulletin 2026: The AI Governance Program Insurers Are Now Asked to Produce
Most AI regulation debates focus on statutes that may or may not pass. Insurance regulators skipped that step. Through a model bulletin adopted state by state, they told carriers that existing unfair-discrimination and trade-practice law already applies to AI, and that examiners will ask to see a written governance program. There is no effective date to plan around — there is a document you either have or do not have when the questionnaire arrives.
Why a Bulletin Instead of a Bill
Insurance is regulated at the state level under a long-established body of law covering unfair trade practices, unfair discrimination, rate adequacy, and claims settlement conduct. Regulators concluded they did not need a new AI statute to reach AI-driven decisions, because those decisions are already the regulated activity: setting a price, declining a risk, denying a claim. The bulletin's core move is to say that outsourcing the reasoning to a model does not change the insurer's accountability for the outcome.
That posture has the same edge-free quality as other existing-law approaches to AI. There is no covered-entity threshold that excludes a small carrier, no annual filing that discharges the duty, and no grandfather date protecting a model that has been in production for years. What replaces those edges is a reasonableness standard, and reasonableness is proved with documents.
The Five Pillars Examiners Look For
Accountability and governance
A named senior owner for AI risk, documented board or executive-level visibility, and written policies that predate the model going live. An after-the-fact policy written in response to an inquiry is visible as such, because version history and approval dates are part of what gets requested.
A complete model inventory
Every system that scores, ranks, flags, or routes a consumer — including vendor scores, third-party data enrichments, rules engines, and fraud flags nobody internally calls AI. Inventories that cover only the pricing model are the most common gap, and claims-side tools are where complaints actually originate.
Risk tiering proportionate to consumer impact
A marketing lookalike audience and a claims-denial recommendation do not warrant the same controls. Tiering lets you defend lighter governance on low-impact tools precisely because you documented heavier governance on the ones that decide coverage and price.
Testing for unfair discrimination and drift
Pre-deployment validation plus ongoing monitoring, with the results retained. Testing that runs but is never recorded gives you no defensible position, and testing that surfaces a disparity which is then filed away without action establishes knowledge without remediation.
Third-party data and model diligence
Documented diligence on vendors, contractual audit and cooperation rights, and enough understanding of a purchased score to explain to a regulator what it measures. "The vendor validated it" is not an answer an examiner accepts on the insurer's behalf.
The Vendor Squeeze: Why Insurtech Feels This First
No state insurance department licenses a machine-learning startup. Yet the bulletin has changed insurtech sales cycles more than it has changed carrier org charts, because the carrier's diligence duty converts directly into a procurement requirement. A carrier that must be able to explain your score to an examiner cannot buy an unexplainable score.
For vendors selling into regulated carriers, the practical readiness list is short and unforgiving: a model card describing inputs, training population, and known limitations; retained testing results by protected class where lawful to measure; a written change log so a carrier can tell which version priced a specific policy; and contract language accepting audit rights and exam cooperation. Vendors that arrive with these close faster than vendors with better model performance and none of them.
AIS Program Readiness Checklist
- ☐List every model touching marketing, eligibility, underwriting, pricing, claims, fraud, and servicing
- ☐Include purchased scores and third-party data enrichments as separate inventory items with their own owners
- ☐Record for each: intended use, consumer-impact tier, owner, vendor, version, and last validation date
- ☐Define which fairness and stability metrics you run pre-deployment and on what monitoring cadence after
- ☐Confirm what you may lawfully collect or infer to test by group in each state before designing the test
- ☐Retain results with the remediation decision attached — including the reasoned decision not to act
- ☐Issue a standard AI diligence questionnaire and keep completed responses on file per vendor
- ☐Negotiate audit rights, exam cooperation, notice of material model changes, and version retention
- ☐Require documentation sufficient to reconstruct which model version produced a specific consumer decision
- ☐Assemble a single exam binder: policies with approval dates, inventory, tiering rationale, testing evidence, vendor files
- ☐Trace three real consumer decisions end to end as a dry run, including any human override and its basis
- ☐Verify complaint-handling can identify when an automated decision was involved and route it accordingly
Check the consumer-facing half of your exposure
Governance files cover the model. Complaints usually start at the quote form, the claims portal, and the automated notice a consumer could not read or complete. RatedWithAI scans those surfaces for accessibility and compliance gaps — start with a free scan.
Scan Your Product for Free →Frequently Asked Questions
Our models do not use race, gender, or any prohibited classification. Are we fine?
Excluding a prohibited input answers the disparate-treatment question and nothing else. Correlated proxies — geography, occupation, vehicle or property characteristics, purchase channel, credit-derived attributes where permitted — can reproduce a prohibited pattern in the output. The examinable question is whether the rating or claims factor producing a disparity has actuarial support and is permitted in that state, which is an evidence question rather than an input-list question.
How does this relate to Colorado's insurance AI law?
Colorado went further than guidance and imposed statutory testing and reporting duties on insurers using external consumer data and algorithms, starting with life insurance. Treat the bulletin as the baseline governance expectation across adopting states and Colorado as a stricter overlay where you write business. Building the general program first makes the state-specific filings a reporting exercise instead of a rebuild.
Does a human reviewer on top of the model resolve the governance question?
Only if the review is substantive. If adjusters or underwriters accept the recommendation in the overwhelming majority of cases, lack the information to challenge it, or are measured on throughput, the model is deciding and the human is documentation theatre. Track override rates and the reasons given — that data is both your best defence and, if the rate is near zero, the clearest evidence against you.
We are a small carrier with two purchased models. How much program is enough?
Proportionate, but not zero. A short written policy, a two-line inventory, vendor diligence files, retained validation results, and a named owner is a defensible program at that scale. What is not defensible is having no document, because the standard is reasonableness and reasonableness is demonstrated rather than asserted.
What usually goes wrong first in a real inquiry?
Two things. The inventory turns out to be incomplete, so a claims-triage or fraud-flag tool surfaces mid-inquiry that no one had governed. And version history is missing, so the carrier cannot say which model priced the policy or flagged the claim that generated the complaint. Both are cheap to fix in advance and expensive to explain afterwards.