FINANCE, POLICY & MARKETSPublished by Paul Ivinskas
fc.The Financial CurrentDAILY INTELLIGENCEWhat matters across finance
Deep-dive library

Model risk across finance: pricing, liquidity, valuations and SR 26-2

7 min read · estimatedAI-generated analysis · Methodology
Historical version · 3 versions · Publication details

First published . This version published .

Version history

About this historical version

Added a hybrid-workflow inventory matrix, material-change decision records and an outcome-based validation handoff. Preserved the distinction between nonbinding guidance, internal choices and other legal obligations.

Compare with an earlier version →
Related research, policy & entities ↓

At a glance

Excerpts from this version
What it covers
SR 26-2 replaces the prior interagency model-risk framework. The expanded analysis adds a component-level inventory, approval boundaries and an evidence-based response to changes in model use.
What would change the conclusion
Evidence of successful implementation would include fewer low-value review tasks alongside better identification and correction of consequential weaknesses. A future interagency statement bringing generative AI into scope, revised statutory duties or changed production use would require reassessment. As of this review, the correct baseline is SR 26-2, with broader governance applied explicitly to the components outside its formal boundary.Read in context
0% through article

Tap a dotted-underlined term for a definition; terms are highlighted once per section. Use Aa in the navigation for reading preferences.

In this article

Start with the current authority

On April 17, 2026, the Federal Reserve, OCC and FDIC issued revised model-risk management guidance. Federal Reserve letter SR 26-2 expressly supersedes and replaces SR 11-7 from 2011 and SR 21-8 from 2021, the latter addressing BSA/AML systems. A policy that still identifies SR 11-7 as the current interagency framework needs an authority review. The new document is , not a statute or a prescriptive regulation. [1][2]

That distinction has practical consequences. The guidance states that it does not establish enforceable standards and that noncompliance with the guidance itself will not produce supervisory criticism. Violations of law or unsafe or unsound practices arising from inadequate management of model risk remain separate matters. Institutions should identify the legal or risk basis for a control rather than describe every internal preference as a regulatory mandate. [2]

Materiality changes the allocation of effort

The revised framework is expected to be most relevant above $30 billion in assets, while recognizing circumstances in which smaller organizations have significant model exposure or complexity. Size is not a sound substitute for understanding use. A small bank relying heavily on a complex third-party underwriting system may have a more consequential model dependency than a larger bank's low-impact internal forecast. [1][2]

Analytical implication: rank review work by the harm a wrong or misused output can cause, the affected exposure, uncertainty and available safeguards. A model used only to prioritize an analyst's reading is not equivalent to a model automatically assigning consumer credit limits. Equally, apparently modest models can become material when embedded in many products or used to feed capital and decisions.

The definition focuses on complex quantitative methods using statistical, economic or financial theories to produce quantitative estimates. Simple arithmetic and deterministic processes without those theoretical foundations are excluded. This narrows the formal inventory boundary but does not make excluded tools harmless. A spreadsheet that miscalculates refunds still needs accuracy and change controls, even if it is not a model for this guidance. [2]

Separate predictive models from generative agents

The attachment excludes generative and agentic AI from its scope because those technologies are evolving rapidly; traditional quantitative models and non-generative, non-agentic AI remain covered. The agencies point to broader risk management for tools outside scope. The Fed's May 1, 2026 AI speech provides further policy context, but a speech should not be elevated into binding requirements. [2][3]

Recommended architecture: inventory a credit workflow's components separately. A default-probability model, deterministic eligibility rules, document-extraction assistant and agent that drafts a case narrative have different failure modes. Link them into a common decision record so a narrow model definition does not leave the surrounding workflow unowned.

For the predictive component, assess development data, target definition, calibration and use limits. For the generative component, test fabricated facts, prompt injection, source attribution and unauthorized actions. For rules, test precedence, completeness and deployment correctness. These are analytical recommendations for different mechanisms, not a claim that SR 26-2 prescribes a single AI control checklist.

Worked example: the same model can have different risk

Hypothetical comparison: Bank A uses a loss model on a $20 million pilot and requires independent review before each credit decision. Bank B applies the same model automatically to a $2 billion portfolio. If a model error understates expected loss by one percentage point across the relevant exposure, the illustrative error is $200,000 at A and $20 million at B. Actual losses would depend on subsequent behavior, use and portfolio dynamics; the arithmetic is a sensitivity, not a forecast.

The model's code can be identical while exposure and reliance differ substantially. Bank B has a stronger economic reason for rigorous testing, independent challenge, monitoring and rollback capacity. Bank A still needs evidence that manual review actually changes decisions when warranted. A nominal human approval step that simply accepts every output may provide little reduction in risk.

A second comparison concerns shared inputs. Five individually modest models may all depend on the same income feed. A single classification error could distort approvals, line management, collections and loss forecasts together. An inventory that counts models without mapping common dependencies can miss that concentration.

A workable transition plan

Recommended first step: map existing policy provisions to current authority, business risk and internal choice. Preserve useful validation and monitoring practices while removing obsolete claims that a withdrawn letter mandates them. Record changes in governance rationale so later reviewers understand why a control was retained, scaled down or replaced.

Next, reassess materiality and ownership using actual production use. Require an accountable business owner to describe what decisions depend on each model, where overrides occur and how failure would be detected. Establish escalation thresholds linked to loss, consumer outcomes or financial reporting significance rather than a uniform calendar exercise.

For vendor models, negotiate enough evidence and access to test intended use. A proprietary algorithm can still be evaluated through benchmark data, stability, error analysis and outcomes. Where transparency remains limited, narrower use, lower limits or additional review may be more defensible than an unsupported declaration of validation. Outsourcing the model does not outsource the decision to rely on it.

What would change the conclusion

The principal opportunity is more proportionate oversight, with scarce review capacity focused on consequential risks. The principal failure mode is interpreting proportionality as permission to remove controls before understanding exposure. Both overclassification and underclassification have costs: unnecessary documentation can delay beneficial changes, while missing a material dependency can produce widespread errors.

Evidence of successful implementation would include fewer low-value review tasks alongside better identification and correction of consequential weaknesses. A future interagency statement bringing generative AI into scope, revised statutory duties or changed production use would require reassessment. As of this review, the correct baseline is SR 26-2, with broader governance applied explicitly to the components outside its formal boundary.

Inventory the decision chain, not just the statistical model

Recommended implementation maps one consequential decision end to end. A default-risk score may be within the guidance’s quantitative-model scope, while a deterministic policy rule and a generative explanation tool have different governance paths. The attachment expressly distinguishes these boundaries; the whole workflow still needs an accountable owner. [2]

The practical goal is to avoid both an inflated model inventory and an unowned gap between components. Use links among components so reviewers can see shared data, version dependencies and where a human can alter the outcome.

Scroll horizontally to see all columns.

ComponentPrimary failure questionRecommended evidence
Predictive scoreDoes the estimate remain fit for its intended use?Outcomes, calibration and use limits
Deterministic eligibility ruleWas the approved rule executed correctly?Boundary tests and deployment comparison
Generative narrativeDoes it invent or misstate the actual decision?Source-linked tests and human review where needed
Shared data pipelineDid meaning or coverage change?Data lineage, reconciliation and change notices
Decision handoffWas the approved output used as intended?Version-linked decision record and override reasons

A change in use can matter more than a code change

Hypothetical: an unchanged score moves from analyst prioritization to automatic line decreases across a large portfolio. Its statistical behavior may be identical, while reliance and potential customer impact rise. Recommended review therefore considers exposure, automation, population and available safeguards, not only software version.

Create a short change record stating the previous use, proposed use, affected population, evidence gaps and permitted limits. Require the risk acceptance to address those changes explicitly. A previous validation is evidence about its evaluated scope; it is not perpetual approval for every future application. These are suggested internal controls, not a claim that the guidance mandates a particular form.

Make the handoff from validation to operations measurable

Recommended release evidence links an approved artifact and intended use to a deployed version and outcome-monitoring plan. Name who will investigate a threshold breach, what can be restricted and how a fallback is activated. An alert without a decision owner is not a completed control.

Test one known error through that chain: can the team identify affected decisions, reproduce the inputs, restrict the faulty use and verify recovery? Also check whether a shared input affects other models. This exercise complements statistical validation by exposing operational dependence. The guidance’s tailored approach supports judgment; it does not make an unexplained green dashboard sufficient evidence of acceptable risk.

Sources

  1. Federal Reserve, SR 26-2, April 17, 2026Official sourceBack to text: ↑1↑2
  2. Federal Reserve/OCC/FDIC, Supervisory Guidance on Model Risk Management, April 17, 2026Official source · PDFBack to text: ↑1↑2↑3↑4↑5↑6
  3. Federal Reserve, Michelle Bowman, Artificial Intelligence in the Financial System, May 1, 2026Official sourceBack to text: ↑

Flag an error or suggest a correction →Public corrections log →