Understanding Second-Line AI Oversight
For the past year, I've been answering questions from governance leaders about second-line AI oversight. They're trying to determine what their risk, compliance, and internal audit teams should do with AI systems. These questions arise from team meetings, Slack threads, and discussions after board presentations that didn't go well.
This confusion is understandable. Your first line, the business units and model developers, builds and owns AI systems. The third line, internal audit, provides independent assurance. But the second line, risk management, compliance, and legal, sits in an awkward middle position. You're supposed to provide oversight without doing the work, challenge without blocking, and scale across numerous models without becoming a bottleneck.
Here's what people are asking and what I've seen work in practice.
Q1: What Should the Second Line Do Differently from Model Developers?
Your second line shouldn't duplicate the first line's work. Model developers handle validation activities, document model limitations, and monitor performance. That's their role.
Your second line sets the standards those teams follow, reviews whether they're meeting them, and escalates when they're not. Specifically:
- Write policies that define acceptable model risk, required validation evidence, and approval thresholds.
- Review validation evidence for completeness and rigor, not to re-run every test.
- Challenge risk assessments when a team rates a high-stakes credit model as low risk.
- Track enterprise-wide AI risk exposure across all business units.
- Escalate to senior leadership when risk appetite is exceeded.
Think of it this way: the first line runs the tests and writes the validation report. The second line reads that report, asks hard questions about methodology gaps, and decides whether it's sufficient to recommend approval.
Q2: How Can We Provide Oversight with Limited Resources?
You can't review everything equally, and you shouldn't try. Risk tiering is your friend.
Start by classifying models using the AI RMF Profile approach or your own risk criteria. Organizations typically tier based on:
- Regulatory classification (is it high-risk under the EU AI Act?)
- Decision impact (does it affect credit, employment, benefits, safety?)
- Data sensitivity (personal data, protected characteristics?)
- Deployment scale (100 users or 10 million?)
Your highest-risk tier might get quarterly second-line review, detailed challenge sessions, and board reporting. Your lowest tier might get annual spot checks and automated monitoring alerts. The middle tier gets something in between.
Document your tiering criteria clearly. When someone asks why their chatbot got more scrutiny than the marketing recommender, you need a defensible answer that isn't "because we felt like it."
Q3: How Do We Avoid Being a Blocker for Developers?
This tension is real, but you can reduce friction.
First, be clear about what triggers second-line involvement. Don't require a full risk review for every model retrain or minor parameter adjustment. Define materiality thresholds: new model deployments, major architecture changes, expansion to new use cases, or material performance degradation.
Second, standardize your review process. Create a review template with specific questions you'll ask every time. Publish expected turnaround times. When developers know what you need and when they'll hear back, they can plan accordingly.
Third, embed review touchpoints into the AI lifecycle rather than bolting them on at the end. A 30-minute risk discussion during model design is more valuable than a two-week review standoff before production deployment.
Finally, measure your own performance. Track how long reviews take, how often you send work back for gaps, and whether you're catching issues that matter. If you're consistently taking three weeks to review low-risk models, you're the blocker.
Q4: What Technical Depth Do Second-Line Reviewers Need?
Your second-line reviewers don't need to write production code, but they can't be completely non-technical either.
You need enough depth to:
- Read and critique a model validation report.
- Understand common AI risks (bias, overfitting, data leakage, adversarial vulnerabilities).
- Ask informed questions about training data quality, test set construction, and performance metrics.
- Recognize when validation evidence is missing or insufficient.
- Interpret model monitoring dashboards and know when drift matters.
You don't need to implement gradient descent or fine-tune a transformer. But if someone tells you their model has 95% accuracy and you can't ask "on what population?" or "compared to what baseline?", you're not providing effective oversight.
Consider building a mixed team: some people with deep technical backgrounds (ex-data scientists, ML engineers) who can engage on methodology, and others with strong risk or compliance expertise who understand regulatory requirements and control frameworks.
Q5: How Do We Handle Vendor Models?
Vendor model risk is one of the hardest second-line challenges, especially with Foundation Model Providers who won't share training data or architecture details.
Your leverage points:
- Vendor Due Diligence before procurement. Require vendors to provide validation evidence, third-party audit reports, and documentation of their own model risk management practices. If they can't or won't, that's a risk factor in your buy/build decision.
- Contractual rights. Negotiate for ongoing performance reporting, incident notification, and the right to conduct or commission independent validation. Some vendors will agree to this; others won't.
- Validation of your implementation, even if you can't validate the underlying model. You can test for bias in your use case, with your data, in your deployment context. You can monitor outputs for quality degradation. You can implement guardrails and human oversight.
- Concentration risk tracking. If you're using the same vendor model across ten different business processes, you've got correlated risk exposure. Track and report this at the enterprise level.
Document what you can't validate and why. When the auditor or regulator asks about your vendor model controls, "we asked but they said no" supported by contract negotiation records is better than silence.
Q6: Who Should Own the AI Model Inventory?
First line should own and maintain it. Second line should verify it's complete and accurate.
Model inventory is operational data. The teams building and deploying models know what exists, where it's running, and what it does. They should be logging this information as part of normal model provisioning and lifecycle management.
Your second line role is to:
- Define what must be tracked (model purpose, risk tier, validation status, approval date, owner, data sources, deployment locations).
- Audit the inventory periodically to find gaps (shadow AI, retired models still running, missing risk classifications).
- Report inventory metrics to leadership (total models, high-risk count, overdue validations, vendor dependencies).
- Challenge completeness when you discover models through other channels that aren't in the inventory.
If second line owns the inventory, you become a data entry service for the first line. That's not oversight; that's administrative burden in the wrong place.
Next Steps for Building Second-Line AI Oversight
If you're building out second-line AI oversight:
- Start with SR 11-7 even if you're not in financial services. It's the clearest articulation of the three lines model for model risk, and most of it translates directly to AI systems.
- Review ISO/IEC 42001 Annex A controls for the organizational structure and accountability requirements that support effective second-line functions. Controls A.6 (Accountability) and A.7 (Risk Management) are particularly relevant.
- Look at the NIST AI RMF Govern function for guidance on organizational roles, responsibilities, and risk culture that enable second-line effectiveness.
And talk to your peers. The organizations getting this right didn't figure it out from standards alone. They iterated, learned from what didn't work, and adapted their approach to their specific risk profile and organizational culture. Your second line won't look exactly like anyone else's, and that's fine.



