Understanding the Source of the Questions
When Cohere, OpenAI, and AI21 Labs released their preliminary practices for deploying large language models, your model risk team likely had immediate questions. The guidance is high-level, leaving you to translate it into governance controls that align with SR 11-7, ISO/IEC 42001, or your internal framework.
These questions arise from real discussions with model risk managers, AI governance leads, and compliance teams who need to understand what these practices mean for validation, third-party risk assessments, and audit readiness.
Q1: Do These Practices Meet SR 11-7's Effective Challenge Requirement?
No, they don't. SR 11-7 requires independent validation that challenges model developers. Practices from the same vendors who build the models don't count as independent review.
You can reference these practices as part of your vendor due diligence under SR 11-7's Vendor Model Risk guidance, but you still need to:
- Conduct your own model risk tiering based on materiality and complexity.
- Document model limitations and use restrictions specific to your use case.
- Establish ongoing monitoring metrics that reflect your risk appetite.
- Maintain validation evidence that demonstrates independent assessment.
Think of vendor practices as one input to your Model Risk Management framework, not a substitute for it.
Q2: Do These Practices Make Us Compliant with the EU AI Act?
No. The EU AI Act sets legal obligations for high-risk AI systems and General-Purpose AI Models, including specific Technical Documentation (Annex IV) requirements, conformity assessment procedures, and Post-Market Monitoring obligations. Voluntary practices from model providers don't create or satisfy legal compliance.
If you're deploying foundation models in EU markets, you need to:
- Determine whether your use case qualifies as high-risk under Annex III.
- Assess whether the underlying model meets the General-Purpose AI Model with Systemic Risk threshold.
- Implement required transparency measures, including Disclosure of AI Interaction where applicable.
- Establish Data Protection Impact Assessment processes if you're processing personal data.
These practices might inform how you operationalize some requirements, but they don't replace your legal compliance analysis.
Q3: Should These Practices Be Included in Our ISO/IEC 42001 Annex A Controls?
Selectively, yes. ISO/IEC 42001 requires you to define controls based on your risk assessment, not adopt every industry recommendation wholesale.
Review the practices against your existing Annex A Controls and risk treatment plan. Where they address gaps in your current controls, incorporate them. Where they conflict with your risk appetite or operational constraints, document why you're taking a different approach.
The standard's Plan-Do-Check-Act (PDCA) cycle means you'll revisit these decisions during management review. Don't treat preliminary vendor guidance as immutable requirements.
Q4: How Do We Validate Models When Providers Won't Share Technical Details?
This is a core challenge with Outsourced Models. You have three practical options:
Option 1: Behavioral validation. Test the model's outputs against your use case requirements. Document performance on your specific tasks, measure bias across relevant demographic segments, and establish baseline error rates. This doesn't tell you how the model works, but it shows whether it works for your application.
Option 2: Enhanced vendor due diligence. Request System Cards or Model Cards that provide more technical detail than the general practices. Ask for specifics on training data characteristics, known limitations, and recommended use restrictions. If the vendor won't provide this, that's risk information.
Option 3: Risk-appropriate deployment. If you can't get sufficient validation evidence, limit the model to lower-risk applications where the impact of failures is manageable. Your risk tiering should reflect validation gaps.
Q5: What Monitoring Metrics Should We Track?
The practices are vague because appropriate monitoring depends on your use case. Start with metrics that map to your risk tiering:
For accuracy and reliability: track prediction error rates, confidence score distributions, and output consistency over time. Set thresholds that trigger model recalibration when performance degrades.
For fairness: measure performance disparities across protected classes relevant to your application. This requires demographic data, which means you need a Data Protection Impact Assessment if you're subject to GDPR.
For security: monitor for adversarial inputs, unusual query patterns that might indicate Adversarial Simulation attempts, and rate limiting violations.
For drift: compare input distributions against your training baseline. Significant distribution shift indicates you're using the model outside its validated operating range.
Q6: What Should We Include in Our Vendor Risk Assessment?
Don't just check a box that says "follows industry practices." Your Vendor Due Diligence needs specific evidence:
- Request documentation of how they implement each practice relevant to your use case.
- Ask for their Post-Market Monitoring approach and how they'll notify you of issues.
- Clarify their Responsible Disclosure process if vulnerabilities are discovered.
- Understand their model versioning and how updates affect your deployment.
- Get specifics on Instructions for Use and documented Model Limitations and Use Restrictions.
If the vendor can't or won't provide this detail, that's a finding in your vendor risk assessment. It might not disqualify them, but it should inform your risk treatment and ongoing monitoring intensity.
Q7: How Do We Prioritize Implementing These Practices?
Map the practices to your existing risk controls and identify gaps. Prioritize based on:
Regulatory requirements first. If a practice addresses a specific EU AI Act obligation or SR 11-7 expectation, it moves up the queue.
Material risks second. Focus on practices that mitigate risks you've identified as material in your AI System Impact Assessment or risk tiering.
Operational feasibility third. Some practices require infrastructure or capabilities you don't have yet. Document these as longer-term improvements in your risk treatment plan.
Your ISO/IEC 42001 management review should track progress against this prioritization. Don't let "preliminary best practices" become a compliance checkbox that distracts from your actual risk management work.
Where to Go for More
The NIST AI RMF Playbook provides practical guidance on operationalizing risk management for AI systems, including foundation models. ISO/IEC 23894 offers structured approaches to AI risk management that complement vendor practices. For financial services specifically, review the model risk management expectations in SR 11-7 and map vendor practices to those requirements explicitly.
Your governance framework should integrate vendor guidance, regulatory requirements, and organizational risk appetite into a coherent approach. Preliminary practices from model providers are a starting point, not a destination.



