A debate is quietly dividing AI governance teams: When you make claims about your AI systems' safety, fairness, or privacy, who should verify those claims? Your internal validation team, or an independent third party?
This isn't academic. The EU AI Act requires Technical Documentation (Annex IV) that substantiates system claims. The NIST AI RMF calls for measurement and evidence throughout the Measure function. ISO/IEC 42001 demands objective evidence for your AI Management System controls. But none of these frameworks explicitly mandate external verification for most use cases.
So where does that leave you?
The Question at Hand
Your team has built an AI system. You've documented its performance metrics, bias testing results, and security controls. You're ready to assert that the system meets your risk thresholds. The question: Is your own validation evidence sufficient, or should you bring in external verifiers to substantiate those claims?
A recent multi-stakeholder report by 58 co-authors from 30 organizations, including the Centre for the Future of Intelligence and Mila, outlines 10 mechanisms to improve the verifiability of AI system claims. These tools range from audit trails to cryptographic proofs, all designed to help developers provide evidence and help stakeholders evaluate that evidence.
But the report doesn't answer the governance question: Who should actually perform that verification?
The Case for Internal Verification
Many practitioners argue that your own model risk team is best positioned to verify AI system claims. They know your systems, your data pipelines, and your organizational context. They can validate continuously rather than waiting for quarterly external audits.
Internal verification fits the SR 11-7 model. Under that guidance, you maintain an independent model validation function that reports outside the development chain. That function reviews model development, tests performance and limitations, and produces Validation Evidence before deployment. It's independent enough to challenge the developers, but embedded enough to move quickly.
This approach scales. When you're deploying dozens of models across business units, external verification for each one becomes a bottleneck. Your internal validators can prioritize based on your Risk Tiering framework, focusing deep reviews on high-risk systems while lighter checks cover lower-tier models.
Cost matters too. External verification isn't cheap, especially when you need domain expertise in both AI and your specific application area. For systems that don't process personal data, don't make high-stakes decisions, and aren't subject to sector-specific regulations, internal validation may be proportionate to the actual risk.
The Case for External Verification
The counterargument: Internal validation has an inherent credibility problem. You're asking stakeholders to trust that your team objectively evaluated your own work. When claims involve safety, fairness, or privacy properties that affect external parties, self-certification doesn't build confidence.
External verifiers bring fresh perspectives. They're not anchored to your design assumptions or organizational incentives. They can identify blind spots your internal team missed, particularly around Contextual Risk Factors that emerge when systems interact with real-world populations and environments.
Regulatory pressure is mounting. While the EU AI Act doesn't mandate third-party conformity assessment for most high-risk systems (it relies on internal checks with notified body oversight for specific annexes), regulators increasingly expect independent validation for material claims. A Data Protection Impact Assessment under GDPR often benefits from external review. Human Rights Due Diligence under frameworks like the UNESCO Recommendation on the Ethics of Artificial Intelligence calls for stakeholder input that goes beyond internal assessment.
For General-Purpose AI Models with Systemic Risk, external scrutiny isn't optional. The General-Purpose AI Code of Practice will likely require independent evaluation of systemic risk claims. Foundation Model Providers can't credibly self-assess whether their models pose economy-wide risks.
Where Practitioners Actually Land
In practice, most organizations use a hybrid model. They maintain internal validation as the primary control, then bring in external verifiers for specific triggers:
Regulatory requirements. If a sector regulator or conformity assessment body requires external validation, you don't have a choice. Financial services firms routinely use external validators for SR 11-7 compliance.
High-stakes systems. When your AI system makes decisions about credit, employment, healthcare, or criminal justice, external verification adds credibility that internal validation can't provide. The reputational cost of getting it wrong exceeds the audit cost.
Novel techniques. When you're deploying Federated Learning, Differential Privacy, or Homomorphic Encryption for the first time, external experts can validate your implementation in ways your internal team may not have the specialized knowledge to do.
Vendor Due Diligence. When you're procuring Outsourced Models, you can't perform internal validation the same way you would for models you built. You need external verification of the Foundation Model Provider's claims, or at minimum, independent review of their System Cards and Model Cards.
Post-incident response. After an AI system causes harm or near-miss, external Root Cause Analysis restores stakeholder trust more effectively than internal review.
Our Take
Mandatory third-party verification for all AI systems is impractical and inefficient. But purely internal validation for high-risk systems is increasingly untenable.
The right answer depends on Materiality. Ask: If this system's claims turn out to be wrong, who bears the cost? If the answer is "primarily us," internal validation may suffice. If the answer is "our customers, the public, or vulnerable populations," you need external verification.
Build your verification strategy around your Risk Tiering framework. Low-risk systems get internal validation with documented evidence. High-risk systems get external validation at critical gates: pre-deployment, after significant Model Recalibration, and during Post-Market Surveillance reviews. Systems in between get periodic external spot-checks.
The 10 verifiability mechanisms in the multi-stakeholder report give you the tools. Audit trails, reproducibility protocols, and structured testing frameworks make both internal and external verification more rigorous. The governance question is who applies those tools, and when.
Don't treat verification as a binary choice. Treat it as a control you calibrate to the risk you're taking and the trust you need to earn.



