October 6, 2026
After SR 26-2: A Control Framework for the Generative and Agentic AI Your MRM Program No Longer Covers

Contents
Key Takeaways
- SR 26-2 replaced SR 11-7 on April 17, 2026. The Federal Reserve, OCC and FDIC issued it jointly, and the OCC published it as Bulletin 2026-13. It also supersedes SR 21-8 and is aimed mainly at banking organizations with more than $30 billion in total assets.
- Generative and agentic AI are outside SR 26-2’s scope. Traditional statistical models and AI that is neither generative nor agentic, such as a gradient-boosted fraud model, remain in scope.
- Supervisory exposure remains. Violations of law under ECOA, BSA/AML, UDAAP and privacy rules, and unsafe or unsound practices, can still draw supervisory action, and third-party risk guidance still applies to AI vendors.
- No US banking agency has issued AI-specific rules for agents yet. The interagency AI request for information announced in April had not been published as of early October 2026, while the EU and Singapore are moving to cover AI agents directly.
- Banks can close the gap with seven controls built on SR 26-2’s own principles: inventory and scoping, materiality tiering, validation and effective challenge, ongoing monitoring, third-party oversight, change control, and board reporting.
- Each control should name an accountable owner, a test and an examiner-ready artifact, mapped to NIST AI RMF and ISO/IEC 42001 so the program holds whatever the RFI concludes.
SR 26-2 is the revised interagency guidance on model risk management that the Federal Reserve, the OCC and the FDIC issued on April 17, 2026. It replaces SR 11-7. In a single footnote, it also places generative and agentic AI outside its scope and tells banks to govern those systems under their existing risk management and governance practices.
That footnote moves a large share of new AI work out of the model risk program without moving it out of the exam. Credit memo drafting, AML narrative generation, KYC document extraction and customer-facing assistants still carry the financial, compliance and conduct risk that model risk management exists to contain. The guidance no longer supplies a method for governing them.
This article covers what changed, which systems stay in scope, and where AI agent regulation in banking stands more broadly. It then sets out a seven-control framework for the systems the agencies left uncovered. Each control names an accountable owner, the test that evidences it, and the artifact an examiner or internal audit will ask to see.
What Replaced SR 11-7? SR 26-2 and OCC Bulletin 2026-13
SR 11-7 was replaced by SR 26-2, the Revised Guidance on Model Risk Management issued jointly by the Federal Reserve, the OCC and the FDIC on April 17, 2026. The OCC published the same guidance as OCC Bulletin 2026-13 and the FDIC as FIL-15-2026. It also supersedes SR 21-8, the 2021 interagency statement on models supporting BSA/AML compliance.
On the OCC side, the announcement rescinded OCC Bulletin 2011-12, OCC Bulletin 2021-19 and OCC Bulletin 1997-24 on credit scoring models. The OCC also withdrew the Model Risk Management booklet of the Comptroller’s Handbook, which held the examination procedures built on the 2011 guidance.
The core architecture survives: development and use, validation, ongoing monitoring, governance, and vendor oversight. The emphasis moves from uniform rigor to rigor proportionate to materiality.
Does SR 26-2 Apply to Generative AI?
No. Footnote 3 of SR 26-2 states that generative AI and agentic AI models are “novel and rapidly evolving” and so are not within the scope of the guidance. The same footnote says a banking organization’s own risk management and governance practices should determine the controls for any system the guidance does not cover.
The footnote also draws the other edge of the boundary. SR 26-2’s principles apply to traditional statistical and quantitative models and to AI models that are neither generative nor agentic. A gradient-boosted fraud model or a machine-learning credit score stays inside model risk management. An LLM drafting a suspicious activity report narrative, or an agent that requests documents and opens cases, does not.
| System type | Banking example | SR 26-2 status | Governed under |
| Traditional statistical model | PD and LGD models, CECL, stress testing, interest rate risk | In scope | Model risk management |
| Non-generative, non-agentic AI | Gradient-boosted fraud detection, ML credit scoring, statistical AML alert scoring | In scope | Model risk management |
| Deterministic rules and arithmetic | Fixed-threshold alert scenarios, spreadsheet calculations | Outside the model definition, unless statistical analysis underpins the design | Operational, IT and compliance controls |
| Generative AI | Credit memo drafting, SAR narrative drafting, retrieval-augmented answers over policy documents | Out of scope (footnote 3) | Bank’s AI governance framework |
| Agentic AI | An agent that gathers KYC documents, calls internal tools, updates case records or messages customers | Out of scope (footnote 3) | Bank’s AI governance framework |
| Hybrid pipeline | An LLM extracts income fields from tax returns that feed an in-scope credit model | Split | Both, with the hand-off documented |
Hybrid pipelines need the most care. SR 26-2 lists the quality of a model’s inputs as a driver of inherent risk. When a generative system produces those inputs, the downstream model owner inherits its error rate whether or not the generative component sits in the MRM inventory.
Record every scoping decision, with the rationale and the approver. In an exam, a documented classification is evidence. An undocumented one reads as a gap.
Why the SR 26-2 Carve-Out Is a Governance Gap, Not an Exemption
The carve-out narrows a guidance document and leaves the bank’s legal and supervisory obligations where they were. SR 26-2 says it sets no enforceable standards, and the same footnote reminds banks that supervisory action can still follow violations of law or unsafe or unsound practices.
That second hook is changing too. The OCC and FDIC adopted a final rule, effective November 2, 2026, that defines an unsafe or unsound practice for the first time. The conduct must break generally accepted standards of prudent operation and have caused, or be likely to cause, material harm to the bank’s financial condition. Matters requiring attention can issue at a lower “could reasonably be expected” threshold. The Federal Reserve did not join the rule.
For OCC- and FDIC-supervised banks using generative and agentic AI, that leaves two main channels for formal criticism:
- Violations of law. ECOA and Regulation B still require specific reasons for adverse action, whatever drafted the notice. The bank still owns every suspicious activity report, including a narrative an LLM wrote. UDAAP and privacy rules still apply to what a customer-facing assistant says and what data it sends out.
- Material financial harm. An agent with write access to case records or payment workflows can produce losses that clear the materiality bar quickly.
“Generally accepted standards of prudent operation” also needs a referent when the model risk guidance is silent. For AI, the most widely adopted ones are NIST AI RMF and ISO/IEC 42001. A program mapped to both gives the bank a defensible answer to what standard it applied.
Vendor oversight has not paused either. The 2023 interagency third-party guidance remains in force while the agencies consult on a principles-based replacement proposed on September 11, 2026 (OCC Bulletin 2026-46), with comments due in mid-November. Both versions keep the bank fully responsible for risk it sources from vendors, including foundation model providers.
The practical failure is an orphaned system. Many bank AI policies route every AI use case through model risk intake. When MRM policies are rewritten to match SR 26-2’s narrower definition, generative and agentic systems can lose their only control owner. The American Bankers Association welcomed the carve-out as regulatory clarity. Clarity about scope still leaves the bank to define the controls.
Where AI Agent Regulation in Banking Stands in October 2026
No US banking agency has issued AI-specific rules for generative or agentic systems. US obligations come through existing law and general guidance, while Singapore and the EU are moving to cover AI agents directly. A bank with an international footprint should build to the broadest scope it faces.
| Jurisdiction | Instrument | Status, October 2026 | What it means for GenAI and agents |
| US federal banking agencies | SR 26-2 / OCC Bulletin 2026-13 | In effect since April 17, 2026; AI RFI announced, not yet published | Out of scope; governed under the bank’s own practices |
| US (OCC, FDIC) | Unsafe or unsound practice rule | Effective November 2, 2026 | Formal criticism tied to material financial harm and violations of law |
| US federal banking agencies and NCUA | Third-party risk management guidance | 2023 guidance in force; replacement proposed September 11, 2026 | Vendor and foundation-model relationships stay in TPRM scope |
| US Treasury | Financial Services AI Risk Management Framework | Released February 19, 2026; voluntary | 230 control objectives aligned to NIST AI RMF, scaled by AI adoption stage |
| US federal and state | Executive Order 14365 | Signed December 11, 2025; preemption proposed, not enacted | State AI laws still apply; DOJ task force can challenge them |
| Colorado | SB 26-189 | Signed May 14, 2026; effective January 1, 2027 | Replaces the Colorado AI Act with narrower duties for automated decision-making in consequential decisions |
| European Union | AI Act, as amended by the Digital Omnibus | Article 50 transparency duties apply from August 2, 2026; Annex III high-risk duties from December 2, 2027 | Chatbot AI-disclosure duties apply now; creditworthiness assessment is high-risk from December 2027 |
| United Kingdom | PRA SS1/23 | In effect since May 2024 for banks with internal-model approval | Technology-agnostic model risk principles with no generative AI carve-out |
| Singapore | MAS Guidelines on AI Risk Management | Consulted November 2025; final text pending | MAS confirmed in August 2026 that agentic AI is in scope |
The asymmetry matters for scoping. SR 26-2 takes agents out of model risk guidance, while MAS writes them in and the EU classifies credit use cases as high-risk regardless of architecture. Treasury’s framework is the closest thing to a US control catalog for these systems, and the framework below cross-references it through NIST AI RMF.
How Should Banks Govern Agentic AI Outside Model Risk Management?
Govern generative and agentic AI through a parallel framework that reuses SR 26-2’s own principles of materiality, effective challenge and ongoing monitoring, adapted to systems that produce text and take actions. The framework below has seven controls. Each names an accountable owner, the test that evidences it, and the artifact an examiner or internal audit will ask for.
Keep the model risk program’s vocabulary where it still fits. Using the same tier scale, the same independence standard for validators and the same issue-tracking workflow lets the board see one picture of aggregate risk instead of two incompatible ones.

| Control | SR 26-2 principle it extends | NIST AI RMF | ISO/IEC 42001 |
| 1. Inventory and scoping | Model inventory; aggregate risk | GOVERN 1.6; MAP 1 | Clause 4.3; Annex A.4 |
| 2. Materiality tiering | Inherent risk, exposure, purpose and use | MAP 5.1 | Clauses 6.1.2 and 6.1.4; Annex A.5 |
| 3. Validation and effective challenge | Conceptual soundness; outcomes analysis; effective challenge | MEASURE 1.3; MEASURE 2 | Annex A.6.2.4 |
| 4. Ongoing monitoring | Ongoing model monitoring | MEASURE 2.4; MANAGE 4.1 | Clause 9.1; Annex A.6.2.6 and A.6.2.8 |
| 5. Third-party and vendor oversight | Vendor and other third-party products | GOVERN 6.1; MANAGE 3.1 and 3.2 | Annex A.10 |
| 6. Change control | Use beyond intended purpose; validation before first use | GOVERN 1.4; MANAGE 2.4 | Clauses 6.3 and 8.1 |
| 7. Board-level reporting | Governance and controls; aggregate risk; role of internal audit | GOVERN 2.3; GOVERN 1.5 | Clauses 5.1, 9.2 and 9.3 |
Treasury’s Financial Services AI Risk Management Framework is built on the NIST AI RMF functions, so the NIST column also points into its control objectives. Its inventory objective, for example, sits under GV-1.6. Mapping to all three keeps the framework valid whatever the agencies’ RFI concludes.
Control 1: AI Inventory and Scoping
Inventory each deployed system rather than each model. An AI model inventory keyed to model IDs misses most generative AI, which arrives as a vendor feature switched on in an existing platform, a developer’s API key, an internal copilot or an agent built inside a business line. One record per deployed use case captures what an examiner will be likely to ask about.
For a generative or agentic system, the record should hold:
- the foundation model, provider and pinned version;
- the system prompt, retrieval sources and guardrail configuration;
- for agents, every tool the system can call, whether each is read or write, and any transaction or spend limits;
- the data classes it touches, including customer and confidential supervisory information;
- where a human reviews output before it takes effect;
- the SR 26-2 scoping decision (MRM, AI governance, or both for hybrid pipelines), with rationale and approval.
Discovery has to look beyond self-reporting. Reconcile the inventory against procurement intake, SSO and API gateway logs, network egress to model provider endpoints, API spend in expense data, and vendor release notes announcing AI features.
- Owner: the business owner of each system attests to its record; the AI governance office (or MRM, if it administers both inventories) maintains the register.
- Test: a quarterly reconciliation of the inventory against discovery sources, plus an internal audit sample traced from egress logs and contracts back to inventory records.
- Artifact: the AI system inventory with scoping classifications, and the reconciliation report listing exceptions and closure dates.
Control 2: Materiality Tiering When the Output Is Text or an Action
SR 26-2 defines materiality as exposure plus purpose, and offers portfolio size as a way to measure exposure. That works for a score. It does not work for a drafted narrative or an agent that files a case, so the exposure dimension needs new measures.
Score exposure on five factors:
- Decision proximity: does the output inform a person, recommend an action, or execute one?
- Reversibility: can the effect be undone before it reaches a customer, a regulator or the ledger?
- Record status: does the output become a regulatory record or customer communication, such as a SAR narrative, an adverse action notice or a disclosure?
- Reach: how many customers, accounts or transactions it touches per day.
- Autonomy: whether a qualified human approves each output before it takes effect.
Purpose keeps the meaning SR 26-2 gives it. Systems used for regulatory compliance or financial risk management rank higher by default.
| Tier | Criteria | Banking example | Minimum controls |
| 1 | Output becomes a regulatory record or customer-facing commitment, or an agent can write to systems of record without pre-approval | SAR narrative drafting; customer assistant that can change account settings | Full independent validation and red-teaming before launch; continuous monitoring; quarterly board-level reporting |
| 2 | Output informs a regulated decision, with mandatory human review | Credit memo summaries; KYC document extraction reviewed by an analyst | Targeted validation; monthly regression testing; sampled human QA |
| 3 | Internal productivity with no regulated effect | Drafting internal meeting notes; code assistants in non-production environments | Inventory record; acceptable-use controls; triggers that force re-tiering |
Tier 3 mirrors how SR 26-2 treats immaterial models: identify them and watch for the conditions that would make them material. A Tier 3 assistant that starts drafting customer emails has changed tier, and the trigger should catch it.
- Owner: second-line AI risk (or MRM) assigns the tier; the business owner attests to the intended use.
- Test: independent re-performance of tier assignments on a sample, and review of re-tiering triggers fired in the period.
- Artifact: the tiering methodology, the scored rationale for each system, and a tier-change log.
Control 3: Validation and Effective Challenge for Generative AI
SR 26-2 names two validation components that translate directly: conceptual soundness and outcomes analysis. It also says interpretability measures or benchmarking against other models may be more practical than theory for some models, which fits generative systems well.
Conceptual soundness becomes a design review. Is the task suitable for a language model at all? Are the retrieval sources authoritative and current? Are tool permissions the minimum the task needs? Are known limitations written down where users will see them? Benchmark the system against an alternative model and against the human process it supports.
Outcomes analysis becomes evaluation on a versioned test set built from real cases with known answers. For a banking workflow, that set should measure:
- faithfulness to source documents and the rate of unsupported claims;
- omission of required elements, such as the who, what, when, where and why of a SAR narrative;
- escalation and refusal behavior when the system should hand off to a person;
- disparities in tone or content across protected classes, for anything touching credit;
- robustness to prompt injection hidden in uploaded documents or retrieved pages;
- for agents, whether each tool call used the right tool with the right parameters and stayed inside its permissions.
Non-determinism changes the statistics. Run each test case several times and report the distribution, then set acceptance thresholds on the failure rate with a stated confidence level. If a model grades outputs, validate the grader against human labels and report the agreement rate, because the evaluator also needs effective challenge.
SR 26-2 defines effective challenge as review by people with the expertise, independence and organizational standing to force change. For generative AI, that means validators who can red-team, who sit outside the build team, and who can block a launch. Where business need forces use before validation, SR 26-2 points to limits and closer monitoring. In practice, run a capped pilot with 100% human review of outputs.
- Owner: an independent validation function in the second line, with generative AI and red-teaming skills.
- Test: the validator re-performs the evaluation on a held-out test set and runs adversarial testing the build team did not design.
- Artifact: a validation report with the test set version, metric distributions, thresholds, limitations and findings, plus the approval memo and any conditions on use.
Control 4: Ongoing Monitoring
SR 26-2 asks whether a model still performs as expected as products, clients, data and conditions change, with monitoring frequency set by materiality. For a generative system, “as expected” has to be measured on live output, because the inputs shift every day and the provider can change the model underneath you.
Useful production signals for banking workflows include:
- the edit rate on drafts, meaning how much analysts change a generated SAR narrative or credit memo before approving it;
- groundedness scores on a daily sample of outputs checked against their sources;
- escalation, refusal and override rates against their baselines;
- guardrail trigger rates, broken down by rule;
- for agents, tool-call failures, permission-denied events and unusual action sequences;
- changes in input mix, such as a new document type entering a KYC pipeline;
- staleness of the retrieval index against the source policy library.
Re-run the validation test set on a schedule set by tier, and immediately after any change covered by Control 6. Pair it with human sampling, because automated metrics miss tone and context errors that a reviewer catches.
Every threshold needs a pre-agreed response. A breach should open an incident, name who decides between tightening guardrails, rolling back or switching the system off, and set a time limit for that decision.
- Owner: the first-line system owner runs monitoring; second-line AI risk reviews breaches and trends.
- Test: thresholds exist for every Tier 1 and Tier 2 system, and a sample of breaches shows escalation and resolution within the agreed time.
- Artifact: the monitoring plan by tier, periodic monitoring reports, the breach and incident log, and immutable action logs for agents.
Control 5: Third-Party Model Risk Management for AI Vendors
Most generative AI in banks runs on someone else’s model. SR 26-2 expects banks to understand a vendor model’s conceptual soundness, design, development data and performance, and to monitor outcomes on an ongoing basis.
A defensible AI vendor file covers:
- the provider’s system or model cards and published evaluation results;
- independent assurance, such as SOC 2 Type II reports and any ISO/IEC 42001 certification;
- contract terms on data use and retention, training on bank data, subprocessors, data residency, incident notification and audit rights;
- advance notice of model version changes and deprecation dates, routed into change control;
- the bank’s own outcome testing on its own use cases, which carries more weight than any vendor claim.
Two patterns cause most findings. The first is embedded AI: a SaaS vendor ships a generative feature in a routine release and nobody re-assesses the relationship. The second is concentration: many use cases depend on one provider or one model family, which is the shared-dependency risk SR 26-2 asks banks to assess in aggregate.
The agencies’ September 2026 proposal would let banks accept residual third-party risk within their risk appetite. That suits opaque model components, provided the acceptance is documented, approved and revisited.
- Owner: third-party risk management, with AI risk input on due diligence; the business relationship owner keeps the file current.
- Test: a contract clause review on a sample of AI vendors, and confirmation that each vendor model has bank-side outcome testing and that version notices reached change control.
- Artifact: AI vendor due diligence files, a contract clause matrix, a provider concentration map, and residual risk acceptance memos.
Control 6: Change Control Across Prompts, Tools and Model Versions
In a traditional model, a change usually means recalibration or redevelopment. A generative system changes behavior when anyone edits the system prompt, refreshes the retrieval index, adds a tool, widens a permission, adjusts a guardrail or upgrades the underlying model. Some providers update models behind a floating alias, so behavior can change with no action by the bank at all.
Treat every behavior-shaping element as versioned configuration:
- Keep prompts, guardrail rules, tool definitions and orchestration code in source control, with a version recorded for each production release.
- Pin model versions for Tier 1 and Tier 2 systems and track provider deprecation dates.
- Define material changes by tier. Any new tool, any expansion of write access, and any model version change should count as material for Tier 1.
- Run the regression test set before every material change reaches production, and attach the results to the change record.
- Keep a tested rollback path and a kill switch that can disable an agent’s write tools without taking the whole system down.
Adding a tool to an agent deserves the same scrutiny as extending a model to a new use. SR 26-2 treats use beyond the intended purpose as added risk that calls for new analysis and a review of controls, and a new tool changes what the system can do.
- Owner: engineering owns the change record; second-line AI risk decides materiality and signs off on Tier 1 material changes.
- Test: trace a sample of production changes to an approval and regression evidence, and scan Tier 1 configurations for unpinned model aliases.
- Artifact: the change log with prompt, model, tool and index versions, attached regression results, approval records, and evidence of a rollback test.
Control 7: Board-Level Reporting
The board needs one view of AI risk, whichever program governs each system. SR 26-2 asks banks to assess model risk in aggregate, including shared data, assumptions and dependencies. Splitting generative and agentic AI into a separate, thinner report hides exactly the concentrations the board should see.
A quarterly report to the risk committee should cover:
- systems by tier and by governance route (MRM, AI governance, or both), with additions and retirements;
- validation status, separating validated, conditionally approved and overdue systems;
- open findings by severity and age;
- monitoring breaches and incidents, with root causes;
- provider concentration across use cases;
- performance against an AI risk appetite statement, for example “no Tier 1 agent writes to a system of record without pre-approval”;
- regulatory changes, including the RFI and EU AI Act dates.
Internal audit’s role follows SR 26-2’s description. It should test whether the AI governance framework is rigorous and effective, not repeat validation work.
- Owner: an accountable executive, typically the CRO or chief AI officer; the board risk committee provides oversight.
- Test: internal audit reconciles board reporting to the inventory, issue logs and incident records.
- Artifact: risk committee AI reports, minutes that record questions and challenge, the AI risk appetite statement, and internal audit’s reports on the framework.
The Framework Applied to Five Banking Workflows
The seven controls apply everywhere, but each workflow can fail in its own way. The table shows where testing and control effort should concentrate.
| Workflow | What goes wrong | Typical tier | Tests that matter most | Controls that matter most |
| Credit decisioning support: memo drafting, financial spreading, reason language | Misread financials feed the decision; stated reasons drift from the scorecard’s actual reasons, creating ECOA and Regulation B exposure | 2; 1 if it drafts adverse action notices | Extraction accuracy against source documents; consistency of reasons with the in-scope model; disparity testing | Documented hand-off to the in-scope credit model; underwriter approval; edit-rate monitoring |
| AML alert triage and SAR narrative drafting | An agent closes true positives; a narrative omits required elements or states facts not in the case file | 1 | Agent dispositions compared with investigator decisions on the same alerts; element completeness; faithfulness to case data | No automated closure without sampled review; investigator sign-off; immutable action logs |
| KYC document review | Fields misread on unfamiliar document types; forged or manipulated documents accepted; instructions hidden in uploaded files | 2; 1 if straight-through | Field-level accuracy by document type and language; adversarial document sets | Confidence thresholds that route to analysts; input-mix monitoring |
| Retrieval over policy and product knowledge bases | Answers cite superseded policy; staff quote the wrong fee or rate; the system ignores document entitlements | 2 | Groundedness and citation accuracy; retrieval recall; entitlement tests across user roles | Index freshness monitoring; change control on the corpus |
| Customer-facing assistants | Misstated product terms (UDAAP exposure); missed complaints or reports of unauthorized transactions that trigger error-resolution duties; personal data leaked | 1 | Red-teaming; complaint and dispute intent recognition; data leakage tests | Output and action guardrails; fast handoff to a person; complaint monitoring; AI disclosure for EU customers |
The AML row shows the scoping split most clearly. SR 26-2 absorbed the 2021 BSA/AML model statement, so the statistical alert-scoring model sits in model risk management. The generative layer that triages its alerts and drafts the narrative does not, even though the SAR it produces is the bank’s regulatory filing.
What Evidence Do Examiners Expect for AI Systems Outside MRM Scope?
Expect examiners and internal audit to ask for the same categories of evidence they see for in-scope models: an inventory, a risk assessment, validation, monitoring, vendor diligence, change records and board reporting. The content differs. For generative and agentic AI it has to cover tool permissions, prompt and model versions, and evaluation results for non-deterministic output.
The new OCC and FDIC rule requires examiners to rest findings on objective facts. Records the bank creates as it operates are the strongest version of those facts, and records assembled the week before an exam are the weakest.
| Question an examiner asks | Evidence that answers it | Control |
| What AI do you use, and how do you know the list is complete? | AI system inventory and the latest reconciliation report | 1 |
| Why is this system governed outside model risk management? | Scoping decision record with rationale and approver | 1 |
| How did you decide how much scrutiny it gets? | Tiering methodology and the scored rationale for the system | 2 |
| Who approved it, against what standard, and what did they find? | Validation report, approval memo and conditions of use | 3 |
| How do you know it still works? | Monitoring plan, recent monitoring reports and the breach log | 4 |
| What happens when it goes wrong? | Incident records and evidence of a tested rollback or kill switch | 4, 6 |
| What do you know about the vendor’s model? | Vendor due diligence file, contract clause matrix and bank-side outcome tests | 5 |
| What changed since the last review? | Change log with versions and attached regression results | 6 |
| Does the board know about it? | Risk committee reports, minutes and the AI risk appetite statement | 7 |
| Has anyone independently checked the program? | Internal audit report on the AI governance framework | 7 |
| Which standard did you apply? | AI policy mapped to NIST AI RMF, ISO/IEC 42001 and Treasury’s FS AI RMF | All |
Build the pack for one Tier 1 system first, such as SAR narrative drafting, and run a mock exam against it. The gaps that surface there usually repeat across the portfolio.
Preparing for the Interagency AI Request for Information
The agencies said in April that the RFI would come “in the near future.” Nearly six months later, it had not been published. An RFI is a step before a proposal, and a proposal is a step before final guidance, so supervisory expectations written specifically for generative and agentic AI are at least two stages away.
Waiting is the riskier choice. Examinations continue in the meantime, and the 2026 issuances so far (SR 26-2, the unsafe-or-unsound rule and the third-party proposal) share a principles-based, tailored approach. A framework built on SR 26-2’s own principles and mapped to NIST AI RMF and ISO/IEC 42001 should need adjustment rather than replacement when guidance arrives.
Use the interval to prepare:
- Run the seven controls on your Tier 1 systems now, so the bank’s position rests on operating experience rather than policy drafts.
- Track the numbers the agencies are likely to ask about: systems by tier, validation findings, monitoring breaches, and the share of AI delivered by vendors.
- Decide whether to comment when the docket opens. Banks that bring data on what works for non-deterministic systems will shape the definitions everyone else inherits.
- Watch the related dockets, including the third-party proposal, whose comment period closes in mid-November 2026.
How Lumenova AI Supports the Framework
Lumenova AI gives model risk, compliance and AI teams one place to run these seven controls and keep the evidence that proves they ran. The aim is an examiner-ready record for the systems SR 26-2 left uncovered, built without waiting for new guidance.
| Control | Lumenova AI capability | What it produces for an exam |
| Inventory and scoping | AI inventory covering in-house, vendor and embedded AI | A single register with scoping and tier decisions |
| Materiality tiering | AI inventory, with each system’s tier and rationale recorded | Recorded tier rationale and re-tiering history |
| Validation and effective challenge | Pre-deployment evaluations with 200+ built-in metrics across drift, bias, robustness and performance | Repeatable validation runs tied to test set versions |
| Ongoing monitoring | Continuous observability across production systems | A time-stamped evidence trail of performance and breaches |
| Third-party oversight | Evaluations run on vendor models against the bank’s own use cases | Bank-side outcome testing for models the bank cannot inspect |
| Change control | Pre-deployment evaluations re-run after prompt, tool or model version changes | Regression evidence attached to each material change |
| Agent actions | Automated guardrails acting as continuous controls on what agents can do | Logged enforcement of tool and action limits |
| Board reporting | Pre-built compliance modules | Reporting aligned to the frameworks the program cites |
Lumenova AI’s forward-deploy team works alongside the bank’s own staff, so the program does not have to be designed and staffed from scratch before the first Tier 1 system is covered. Book a discovery call to see the platform in action.
Frequently Asked Questions
OCC Bulletin 2026-13 is the OCC’s issuance of the revised interagency model risk management guidance, dated April 17, 2026. It rescinds OCC Bulletins 2011-12, 2021-19 and 1997-24 and announces a planned request for information on banks’ use of AI, including generative and agentic AI.
The guidance says it is most relevant to banking organizations above $30 billion in total assets. It may still be relevant to smaller banks with significant model risk, for example because their models are complex or widespread, or because they run activities outside traditional community banking.
SR 26-2 states that it sets no enforceable standards and that non-compliance with it will not, by itself, lead to supervisory criticism. Supervisory action can still follow violations of law or unsafe or unsound practices caused by weak model risk management. For OCC and FDIC banks, a new rule defining unsafe or unsound practice takes effect November 2, 2026.
Run agents through a parallel AI governance framework built on SR 26-2’s own principles. Inventory each agent with its tools and permissions, tier it by decision proximity and reversibility, validate it with repeated evaluations and red-teaming, monitor its actions, control changes to prompts, tools and model versions, and report to the board.
Expect requests for an AI inventory with scoping decisions, a tiering rationale, a validation report and approval, monitoring reports and breach logs, vendor due diligence, a change log with regression results, and board reporting. Records created during normal operation carry more weight than documents assembled for the exam.
Yes. The 2023 interagency third-party guidance remains in force, and its proposed September 2026 replacement keeps banks fully responsible for risk sourced from vendors. Generative AI delivered by a provider or embedded in a SaaS product sits inside third-party risk management even though it is outside SR 26-2.
Use validators with red-teaming skills who are independent of the build team and have authority to block launch. Test on versioned sets of real cases, run each case several times, set thresholds on failure rates, and validate any model used as a grader against human judgments.