The question seems rhetorical. Of course, banks should validate their models. SR 11-7 requires it, and model risk management isn't optional in financial services. But dig deeper, and you'll find a real debate among model risk teams: Does traditional validation really apply to Generative AI models, especially when they're used for internal productivity rather than customer-facing decisions?
This isn't just semantics. The question highlights a tension between established risk frameworks and emerging technology that doesn't fit neatly into existing categories.
The Case Against Formal Validation
Skeptics offer several practical arguments. First, many banks use GenAI for tasks that don't meet SR 11-7's definition of a "model", tools that help draft emails, summarize documents, or generate code snippets. These applications don't quantify risk, inform business decisions, or measure financial value like credit scorecards or stress testing models do.
Second, the technology itself resists traditional validation methods. You can't validate a model when you don't control its training data, can't inspect its architecture, and receive updates from the Foundation Model Provider without notice. The reproducibility that underpins conventional validation evidence simply doesn't exist. How do you write a validation report for a black box that changes every few months?
Third, there's the resource argument. Model risk teams are already stretched thin validating credit models, market risk models, and anti-money laundering systems. Adding every GenAI chatbot to the validation queue would overwhelm most departments, potentially diverting attention from models that pose genuine financial risk.
Some practitioners also point to the regulatory gap. Unlike credit risk models or fair lending algorithms, GenAI applications don't yet have specific regulatory guidance. The EU AI Act covers certain high-risk AI systems, but many internal GenAI tools fall outside those categories. NIST AI RMF provides a framework but doesn't mandate validation. Why impose validation rigor when regulators haven't explicitly required it?
The Case for Validation (Even Now)
The counterargument starts with a simple premise: SR 11-7 doesn't exempt new technologies just because they're unfamiliar. The guidance defines a model as "a quantitative method, system, or approach that applies statistical, economic, financial, or mathematical theories, techniques, and assumptions to process input data into quantitative estimates." If your GenAI application processes inputs to generate outputs that inform decisions, even internal ones, it's a model.
More importantly, the risks are real. GenAI models can hallucinate confidential information, embed biases in customer communications, generate incorrect financial calculations, or leak proprietary data through prompt injection. A chatbot that helps loan officers draft denial letters is absolutely a compliance risk, regardless of whether it fits traditional model categories.
The validation framework also provides essential controls that GenAI desperately needs. Conceptual soundness review forces you to understand what the model can and can't do reliably. Ongoing monitoring catches performance degradation before it reaches customers. Vendor Due Diligence ensures your Foundation Model Provider has adequate security controls. These aren't bureaucratic exercises; they're the difference between controlled experimentation and reckless deployment.
Consider the alternative: If you skip validation, how do you document model limitations and use restrictions? How do you perform root cause analysis when something goes wrong? How do you demonstrate to examiners that you understood the risks before deployment? The absence of specific GenAI regulations doesn't eliminate the need for risk management, it makes documentation more critical, not less.
Where Practitioners Actually Land
Most model risk teams aren't choosing one extreme or the other. They're building hybrid approaches that acknowledge both the necessity of oversight and the practical constraints of validating foundation models.
The emerging consensus involves risk tiering. Internal productivity tools with human review get lighter-touch validation focused on data security and output monitoring. Customer-facing applications or those that influence material decisions get full validation, even if that means validating the wrapper application and use case rather than the base model itself.
Many teams are also adapting their validation evidence requirements. Instead of demanding full reproducibility, they're documenting model limitations, testing outputs against known scenarios, and establishing ongoing monitoring thresholds. They're treating Foundation Model Provider documentation the way they'd treat vendor model documentation, with healthy skepticism and supplemental testing.
The smartest teams are also preparing for regulatory evolution. The EU AI Act's General-Purpose AI Code of Practice will establish transparency requirements for General-Purpose AI Model providers. Future guidance will likely clarify expectations for financial services. Building validation infrastructure now means you won't be scrambling when requirements formalize.
Our Take
Banks should validate GenAI models, but validation needs to evolve with the technology.
The core principles of SR 11-7 remain sound: understand what you're deploying, document its limitations, monitor its performance, and maintain effective challenge. But the execution must adapt. You can't perform traditional backtesting on a model you didn't train. You can't reproduce results when the Foundation Model Provider updates parameters without notice.
What you can do is validate the application layer. Test how your GenAI tool performs on representative scenarios. Document the use case boundaries and the controls that prevent misuse. Establish monitoring for output quality, bias indicators, and security anomalies. Perform Vendor Due Diligence on your provider's risk management practices. Require Disclosure of AI Interaction when appropriate.
The validation report might look different, more focused on contextual risk factors and operational controls than statistical performance metrics. But the underlying discipline remains essential. Your examiners won't accept "it's GenAI, we can't validate it" any more than they'd accept "it's machine learning, we can't explain it."
The banks getting this right aren't debating whether to validate. They're building validation frameworks flexible enough to cover both traditional statistical models and emerging AI applications, recognizing that model risk management is a principle, not a checklist.



