RatedWithAI

RatedWithAI

Accessibility scanner

Financial RegulationSeptember 12, 2026

SR 11-7 and Generative AI: The 2011 Model Risk Rule Your Bank Deal Runs Into

Every "we're waiting for AI regulation before we deploy" conversation in banking misreads the situation. The binding constraint is fifteen years old, technique-neutral by design, and it has already stopped more AI features than any statute passed since 2024.

Technique-Neutral
SR 11-7 defines a model by function, not by algorithm
Non-Transferable
Using a vendor model does not move the responsibility
Effective Challenge
Undocumented systems give validators nothing to test

A Definition Written Before Transformers, Wide Enough to Hold Them

The Federal Reserve's SR 11-7 and the OCC's parallel Bulletin 2011-12 define a model as a quantitative method that applies statistical, economic, financial, or mathematical theories and assumptions to process input data into quantitative estimates used in decision making. The definition is functional. It does not care whether the method is a logistic regression, a gradient-boosted tree, or a hundred-billion-parameter language model.

This is why "generative AI isn't really a model" arguments fail inside banks. The question a model risk management (MRM) team asks is narrower and harder to dodge: does the output inform a decision? If an LLM summarization step sits upstream of a human credit decision, an alert disposition, or a customer-communication approval, the MRM function will scope it in — as a model, or as a tool subject to equivalent control expectations. Either way, someone has to document it.

The Three Pillars, Applied to an LLM Feature

Conceptual soundness

Why is this approach appropriate for this use? For a generative feature the honest answer is rarely 'the benchmark score was high.' It requires a design rationale tied to the specific task, the data it sees, and the failure modes the task tolerates.

Ongoing monitoring

Performance must be tracked in production, not just at validation. Generative outputs make this harder because there is often no ground-truth label — which means the monitoring plan has to define proxy signals and escalation thresholds up front.

Outcomes analysis

Compare outputs to actual results. Aggregate accuracy is not enough; validators look at segment-level error, because a model that performs well on average and badly on a protected or low-volume segment is the pattern fair-lending exam findings are made of.

Effective challenge

Independent, competent, empowered review. This is the pillar vendors control least and influence most — the quality of what you hand over determines whether the challenge can happen at all.

The Vendor Problem SR 11-7 Anticipated

The guidance addresses third-party models directly, and its position has not softened with time: a bank that uses an externally developed model is still responsible for validating it. The bank must understand the model's components, assumptions, and limitations well enough to challenge them, and must have contingency plans for when the vendor product changes or goes away.

Generative AI adds a wrinkle the 2011 drafters did not face. A traditional vendor model is a versioned artifact that sits still. A feature built on a hosted foundation model can behave differently next Tuesday because the provider shipped an update. From an MRM standpoint that is uncontrolled change to a validated model, and it is the single most common reason a promising AI pilot stalls at the second gate. The vendors who get through are the ones who arrive with a version-pinning story, a change-notification commitment in the contract, and a documented re-testing trigger — before anyone asks.

Vendor Readiness Checklist for Bank AI Sales

1. Scope and Intended Use
  • Write a one-page intended-use statement naming the decisions the output may and may not inform
  • State known limitations explicitly — an empty limitations section reads as an incomplete assessment, not a strong product
  • Identify whether a human reviews each output, samples them, or does neither
2. Evidence Package
  • Document the design rationale for the approach, not just the results
  • Report evaluation results by segment as well as in aggregate, with the error analysis attached
  • Describe the data used for evaluation and how it relates to the bank's population
  • Include a reproducible test the bank can run against its own data
3. Change and Version Control
  • Name every upstream foundation model and whether its version can be pinned
  • Commit contractually to advance notice of material model changes
  • Define the re-testing trigger: what change forces revalidation, and who initiates it
  • Maintain a change log the bank can attach to its model inventory record
4. Monitoring and Exit
  • Specify production monitoring signals, thresholds, and who is notified on breach
  • Provide the data the bank needs for its own ongoing monitoring, not just your dashboard
  • Document a contingency and exit path if the feature is withdrawn or the provider changes terms

Check what your product shows a reviewer

Diligence starts on your website long before the questionnaire. RatedWithAI scans your public surface and shows the trust, disclosure, and accessibility gaps a reviewer will find first.

Scan Your Product for Free →

Frequently Asked Questions

Is SR 11-7 a regulation or guidance?

It is supervisory guidance, not a rule with its own civil penalties. In practice that distinction matters less than it sounds: examiners assess model risk management against it, findings and matters requiring attention are written against it, and a bank's internal policy typically converts it into binding internal requirements. A vendor who argues that guidance is optional is arguing with the wrong party.

Does a retrieval-augmented chatbot for internal knowledge count as a model?

It depends on what the answers are used for. A tool that helps an employee find a policy document is usually handled as a productivity tool with lighter controls. The same tool, once employees rely on its synthesized answer to decide how to treat a customer or dispose of an alert, starts to look like a decision input — and MRM scoping follows use, not the label in the product name.

How does this relate to third-party risk management requirements?

They stack. Interagency third-party risk management expectations govern the vendor relationship — due diligence, contracts, ongoing oversight, termination — while SR 11-7 governs the model itself. An AI vendor selling into a bank should expect both reviews, run by different teams, sometimes asking for the same artifact in different formats.

Do credit unions and non-bank lenders face the same expectations?

Not identically, but the direction is the same. NCUA supervision, state regulators, and the fair-lending framework all push toward documented model governance, and enterprise buyers increasingly apply bank-style diligence regardless of charter. A vendor that builds an SR 11-7-shaped evidence package can reuse most of it across financial-services buyers.

Related Guides