Models are part of everyday financial decisions
A financial model turns assumptions and observations into an estimate used to make a decision. In banking and fintech, that decision may concern product pricing, cash needs, fraud review, portfolio value or staffing—not only who receives a loan. The first question for a reader is what action depends on the output and how an error would change the result.
Analysis: a useful model can make uncertainty visible and improve consistency. It can also encourage false confidence if users overlook missing data or apply it to a different task. Model-risk management is economically useful when it improves those decisions; the formal guidance explains a supervisory framework rather than replacing the business question.
Start with the current authority
On April 17, 2026, the Federal Reserve, OCC and FDIC issued revised model-risk management guidance. Federal Reserve letter SR 26-2 expressly supersedes and replaces SR 11-7 from 2011 and SR 21-8 from 2021, the latter addressing BSA/AML systems. A policy that still identifies SR 11-7 as the current interagency framework needs an authority review. The new document is , not a statute or a prescriptive regulation. [1][2]
That distinction has practical consequences. The guidance states that it does not establish enforceable standards and that noncompliance with the guidance itself will not produce supervisory criticism. Violations of law or unsafe or unsound practices arising from inadequate management of model risk remain separate matters. Institutions should identify the legal or risk basis for a control rather than describe every internal preference as a regulatory mandate. [2]
Materiality changes the allocation of effort
The revised framework is expected to be most relevant above $30 billion in assets, while recognizing circumstances in which smaller organizations have significant model exposure or complexity. Size is not a sound substitute for understanding use. A small bank relying heavily on a complex third-party underwriting system may have a more consequential model dependency than a larger bank's low-impact internal forecast. [1][2]
Analytical implication: rank review work by the harm a wrong or misused output can cause, the affected exposure, uncertainty and available safeguards. A model used only to prioritize an analyst's reading is not equivalent to a model automatically assigning consumer credit limits. Equally, apparently modest models can become material when embedded in many products or used to feed capital and decisions.
The definition focuses on complex quantitative methods using statistical, economic or financial theories to produce quantitative estimates. Simple arithmetic and deterministic processes without those theoretical foundations are excluded. This narrows the formal inventory boundary but does not make excluded tools harmless. A spreadsheet that miscalculates refunds still needs accuracy and change controls, even if it is not a model for this guidance. [2]
Three uses, three different costs of being wrong
These illustrative applications show why a single accuracy score is insufficient. A deposit model estimates how balances and rates may respond to market changes. A valuation model estimates the worth of expected future cash flows. A payment-fraud model helps decide whether to accept, reject or review a transaction. Each supports a different action and has a different error cost.
Hypothetical treasury sensitivity: a bank models $500 million of interest-bearing deposits with a 30% beta after a one-percentage-point benchmark increase. With unchanged average balances for a full year, the projected additional annual interest expense is $1.5 million. At a realized beta of 60%, it is $3 million—a $1.5 million difference before balance migration or hedging. A numerically modest assumption change can therefore matter materially to earnings planning.
Analysis: for valuation, test how plausible repayment timing and discount-rate assumptions change the estimate rather than treating one output as a certain sale price. For payment fraud, weigh losses prevented against legitimate transactions blocked and the cost of review. These uses need outcome measures suited to their purpose. A model can perform well statistically and still be a poor decision aid if the cost of its errors or the population using it has changed.
Separate predictive models from generative agents
The attachment excludes generative and agentic AI from its scope because those technologies are evolving rapidly; traditional quantitative models and non-generative, non-agentic AI remain covered. The agencies point to broader risk management for tools outside scope. The Fed's May 1, 2026 AI speech provides further policy context, but a speech should not be elevated into binding requirements. [2][3]
Recommended architecture: inventory a financial workflow’s components separately. A quantitative forecast, deterministic limit rule, document-extraction assistant and generative explanation can fail in different ways. Map how outputs pass between them so the formal model definition does not leave a consequential business dependency unexamined. The credit workflow below remains one application of that broader principle.
For the predictive component, assess development data, target definition, calibration and use limits. For the generative component, test fabricated facts, prompt injection, source attribution and unauthorized actions. For rules, test precedence, completeness and deployment correctness. These are analytical recommendations for different mechanisms, not a claim that SR 26-2 prescribes a single AI control checklist.
Worked example: the same model can have different risk
Hypothetical comparison: Bank A uses a loss model on a $20 million pilot and requires independent review before each credit decision. Bank B applies the same model automatically to a $2 billion portfolio. If a model error understates expected loss by one percentage point across the relevant exposure, the illustrative error is $200,000 at A and $20 million at B. Actual losses would depend on subsequent behavior, use and portfolio dynamics; the arithmetic is a sensitivity, not a forecast.
The model's code can be identical while exposure and reliance differ substantially. Bank B has a stronger economic reason for rigorous testing, independent challenge, monitoring and rollback capacity. Bank A still needs evidence that manual review actually changes decisions when warranted. A nominal human approval step that simply accepts every output may provide little reduction in risk.
A second comparison concerns shared inputs. Five individually modest models may all depend on the same income feed. A single classification error could distort approvals, line management, collections and loss forecasts together. An inventory that counts models without mapping common dependencies can miss that concentration.
A workable transition plan
Recommended first step: map existing policy provisions to current authority, business risk and internal choice. Preserve useful validation and monitoring practices while removing obsolete claims that a withdrawn letter mandates them. Record changes in governance rationale so later reviewers understand why a control was retained, scaled down or replaced.
Next, reassess materiality and ownership using actual production use. Require an accountable business owner to describe what decisions depend on each model, where overrides occur and how failure would be detected. Establish escalation thresholds linked to loss, consumer outcomes or financial reporting significance rather than a uniform calendar exercise.
For vendor models, negotiate enough evidence and access to test intended use. A proprietary algorithm can still be evaluated through benchmark data, stability, error analysis and outcomes. Where transparency remains limited, narrower use, lower limits or additional review may be more defensible than an unsupported declaration of validation. Outsourcing the model does not outsource the decision to rely on it.
What would change the conclusion
The principal opportunity is more proportionate oversight, with scarce review capacity focused on consequential risks. The principal failure mode is interpreting proportionality as permission to remove controls before understanding exposure. Both overclassification and underclassification have costs: unnecessary documentation can delay beneficial changes, while missing a material dependency can produce widespread errors.
Evidence of successful implementation would include fewer low-value review tasks alongside better identification and correction of consequential weaknesses. A future interagency statement bringing generative AI into scope, revised statutory duties or changed production use would require reassessment. As of this review, the correct baseline is SR 26-2, with broader governance applied explicitly to the components outside its formal boundary.
For business users, the practical evidence is whether forecasts explain important misses, valuations show credible uncertainty, and production decisions improve after identified weaknesses are corrected. Review effort should be judged alongside the decisions it improves. A growing inventory or a faster approval cycle alone demonstrates neither sound models nor useful financial outcomes.
Inventory the decision chain, not just the statistical model
Recommended implementation maps one consequential decision end to end. A default-risk score may be within the guidance’s quantitative-model scope, while a deterministic policy rule and a generative explanation tool have different governance paths. The attachment expressly distinguishes these boundaries; the whole workflow still needs an accountable owner. [2]
The practical goal is to avoid both an inflated model inventory and an unowned gap between components. Use links among components so reviewers can see shared data, version dependencies and where a human can alter the outcome.
Scroll horizontally to see all columns.
| Component | Primary failure question | Recommended evidence |
|---|---|---|
| Predictive score | Does the estimate remain fit for its intended use? | Outcomes, calibration and use limits |
| Deterministic eligibility rule | Was the approved rule executed correctly? | Boundary tests and deployment comparison |
| Generative narrative | Does it invent or misstate the actual decision? | Source-linked tests and human review where needed |
| Shared data pipeline | Did meaning or coverage change? | Data lineage, reconciliation and change notices |
| Decision handoff | Was the approved output used as intended? | Version-linked decision record and override reasons |
A change in use can matter more than a code change
Hypothetical: an unchanged score moves from analyst prioritization to automatic line decreases across a large portfolio. Its statistical behavior may be identical, while reliance and potential customer impact rise. Recommended review therefore considers exposure, automation, population and available safeguards, not only software version.
Create a short change record stating the previous use, proposed use, affected population, evidence gaps and permitted limits. Require the risk acceptance to address those changes explicitly. A previous validation is evidence about its evaluated scope; it is not perpetual approval for every future application. These are suggested internal controls, not a claim that the guidance mandates a particular form.
Make the handoff from validation to operations measurable
Recommended release evidence links an approved artifact and intended use to a deployed version and outcome-monitoring plan. Name who will investigate a threshold breach, what can be restricted and how a fallback is activated. An alert without a decision owner is not a completed control.
Test one known error through that chain: can the team identify affected decisions, reproduce the inputs, restrict the faulty use and verify recovery? Also check whether a shared input affects other models. This exercise complements statistical validation by exposing operational dependence. The guidance’s tailored approach supports judgment; it does not make an unexplained green dashboard sufficient evidence of acceptable risk.
Sources
- Federal Reserve, SR 26-2, April 17, 2026Official sourceBack to text: ↑1↑2↑3
- Federal Reserve/OCC/FDIC, Supervisory Guidance on Model Risk Management, April 17, 2026Official source · PDFBack to text: ↑1↑2↑3↑4↑5↑6↑7
- Federal Reserve, Michelle Bowman, Artificial Intelligence in the Financial System, May 1, 2026Official sourceBack to text: ↑