Skip to main content
Frontier AI Safety Without a PlaybookCompliance & Audit
5 min readFor AI Governance Leaders

Frontier AI Safety Without a Playbook

When a frontier AI model transitions from development to release, who ensures it's safe? The source material highlights a significant gap: while there's plenty of expert commentary on AI policy frameworks, the actual evaluation criteria used by governments to approve models like OpenAI's remain unclear. For AI governance leaders, this lack of transparency poses a practical challenge. You're tasked with building risk management frameworks without knowing what regulators will measure when they arrive.

The Challenge

Consider what's missing. Experts at CSET have analyzed the balance between innovation and regulation in AI policy. The Trump administration issued an AI executive order on restrictions. There's commentary on U.S.-China AI competition and open-access AI datasets. Yet, nowhere in this policy discourse are government evaluation criteria for frontier model safety documented.

This isn't just an academic issue. If your organization deploys or develops large language models, you need to know: What tests did the model pass? What risk thresholds were applied? What documentation was exchanged? Without published criteria, governance teams are left guessing regulatory requirements.

The Environment and Constraints

The regulatory environment for frontier AI models is under pressure. Governments aim to foster innovation while preventing catastrophic risks. They must balance transparency with national security concerns and write rules for systems evolving faster than policy cycles.

For organizations, this creates asymmetric information risk. You don't know if your model validation approach will satisfy future regulatory scrutiny because the government hasn't published its safety evaluation criteria. You can't benchmark against peers because frontier model evaluations are confidential. International standards don't yet specify quantitative thresholds for "frontier" versus "general-purpose" AI models.

The EU AI Act establishes categories like high-risk AI and General-Purpose AI Models, but it doesn't specify the tests a government would run before allowing a GPT-class model to launch. SR 11-7 provides a model risk management structure for financial services, but it predates large language models. ISO/IEC 42001 offers an AI Management System framework, but it's process-oriented, not outcome-specific.

The Approach (That Doesn't Exist Yet)

What should a government evaluation framework include? Without documented precedent, we can infer requirements from first principles and existing regulatory patterns.

Pre-release technical evaluation. Expect adversarial simulation across NIST AI RMF risk categories: harmful content generation, privacy violations, security vulnerabilities, and fairness failures. The evaluation should involve red teaming by independent experts, not just the model developer's team. Results should be quantified with reproducible test sets, not narrative summaries.

Capability benchmarking. Regulators need defined thresholds. At what performance level on what tasks does a model shift from general-purpose to systemic risk? The EU AI Act mentions systemic risk but doesn't operationalize it. A real evaluation framework would specify: "Models exceeding X score on Y benchmark require additional controls Z."

Systemic risk assessment. This extends beyond the model itself. Consider user numbers, deployment scale, affected critical infrastructure, and misuse vectors. The assessment should follow ISO/IEC 42005 impact assessment structure, but with quantitative risk tiering.

Ongoing monitoring commitments. Pre-release evaluation is necessary but insufficient. Approval should require Post-Market Surveillance with defined reporting triggers. If the model's error rate on protected classes exceeds a threshold, if adversarial attacks succeed above a baseline, or if misuse incidents form a pattern, the developer reports back and regulators re-evaluate.

Disclosure requirements. What does the public see? At minimum: a System Card describing capabilities and limitations, aggregated performance metrics on safety benchmarks, and instructions for use specifying prohibited applications. The EU's General-Purpose AI Code of Practice will likely formalize some of this, but it's not available yet.

Results and Metrics (We Don't Have)

The problem is clear: we can't measure what hasn't been implemented. No government has published the evaluation results that led to a frontier model's approval. We don't know if OpenAI's models met specific safety thresholds or if the evaluation was qualitative. We don't know if there were required model modifications before release or if ongoing monitoring revealed post-launch issues.

This absence of metrics creates compliance risk for your team. You're building model validation processes without knowing if they'll satisfy regulatory expectations. You're setting internal risk thresholds without external calibration. You're making deployment decisions in a policy vacuum.

What They Would Do Differently

If regulators were designing this evaluation framework today with hindsight, several changes seem obvious.

Publish the criteria before requiring compliance. Organizations can't meet standards they can't see. The EU AI Act's phased implementation recognizes this, but even there, the General-Purpose AI Code of Practice is still being finalized. Regulators should release evaluation rubrics, test methodologies, and decision thresholds as public documents.

Establish independent evaluation bodies. Allowing model developers to self-certify safety creates conflicts. The evaluation should involve third parties with Adversarial Simulation expertise, domain-specific risk knowledge, and no commercial relationship to the developer.

Require reproducible evidence. Narrative safety claims aren't validation evidence. The evaluation should produce quantitative results on standardized benchmarks that other researchers can verify. This doesn't mean publishing model weights, which creates its own risks, but it does mean publishing enough methodology that the claimed safety properties can be independently assessed.

Build in adaptation mechanisms. Frontier AI capabilities change fast. The evaluation framework needs scheduled review cycles, not static requirements. If a new attack vector emerges or a new capability benchmark becomes standard, the framework should incorporate it within a defined timeline.

Takeaways for Your Team

While you wait for governments to publish evaluation criteria, you can't pause AI governance work. Here's what you can do now.

Build your own evaluation framework using existing standards as scaffolding. Map your model validation process to SR 11-7 if you're in financial services, or to ISO/IEC 23894 for risk management more broadly. Document your adversarial simulation approach. Establish quantitative thresholds for model performance, fairness metrics, and security controls. When regulators do publish requirements, you'll have a structured baseline to compare against.

Treat transparency as a strategic asset, not a compliance burden. If you're deploying frontier models, publish System Cards voluntarily. Document your model's limitations and use restrictions. Create instructions for use that specify prohibited applications. This transparency builds trust and demonstrates governance maturity before regulators require it.

Participate in standard-setting processes. The General-Purpose AI Code of Practice is being developed now. ISO technical committees are writing AI standards. Industry groups are proposing safety benchmarks. Your participation shapes what the eventual requirements will be, and it gives you early visibility into regulatory direction.

Don't wait for perfect information to act on material risks. The absence of government evaluation criteria doesn't mean you lack risk management obligations. If your model can generate harmful content, implement content filtering. If it processes personal data, conduct a Data Protection Impact Assessment. If it makes decisions affecting individuals, establish human oversight. These controls are defensible under existing frameworks even if frontier-specific regulations haven't arrived yet.

The gap in government evaluation criteria for frontier AI safety is real, but it's temporary. When those criteria do emerge, they'll likely draw from the governance practices organizations are building today. Your job is to make sure your practices are defensible, documented, and grounded in recognized standards, so you're ready when the playbook finally gets published.

You Might Also Like