Scope - What This Guide Covers
This guide focuses on the audit and risk management challenges posed by Frontier AI models capable of autonomous vulnerability discovery and exploit development. It's tailored for internal auditors assessing your organization's readiness for AI-enabled cyber threats, whether you're in financial services, energy, transportation, or tech infrastructure.
You'll find:
- Core concepts about Frontier AI offensive capabilities
- Sector-specific requirements
- Steps to implement over the next 90 days
- Common control gaps exposed by AI-enabled threats
- A quick reference table for sector-specific regulatory triggers
Key Concepts and Definitions
Frontier AI Models: Advanced AI models with reasoning abilities to autonomously identify zero-day vulnerabilities and create exploits without human help. Anthropic's Mythos produced 181 working exploits in tests, compared to just two by their previous model.
Multi-Vector Attack: Exploitation across multiple attack surfaces (network, messaging systems, cloud) executed simultaneously. Traditional incident response plans assume sequential attacks.
Memory-Unsafe Programming Languages: Languages like C and C++ that don't prevent buffer overflows and similar vulnerabilities. Your oldest, least-maintained code in these languages is your highest-risk attack surface.
Project Glasswing: Anthropic's initiative restricting Mythos access to about 40 organizations. Equivalent capability is expected to reach adversaries in 6 to 24 months.
Zero-Day Vulnerability: A security flaw unknown to the vendor with no existing patch. Mythos can autonomously identify these across major operating systems and browsers.
Requirements Breakdown
Financial Institutions (NYDFS / GLBA)
If you're under NYDFS cybersecurity regulations (23 NYCRR 500) or GLBA safeguards:
Risk Assessment Update: Frontier AI models change your threat landscape materially. Your risk assessment must reflect this change. NYDFS 500.09 requires annual assessments, but material changes need interim updates.
Incident Response Plan: NYDFS 500.16 and GLBA require plans designed for current threats. Your plan must handle simultaneous multi-vector scenarios, not just sequential attacks.
Board Reporting: NYDFS 500.04 requires your CISO to report material cybersecurity issues to the board. Autonomous exploit capability meets this threshold.
Energy Sector (TSA Security Directives)
For pipeline or critical energy infrastructure operators:
OT Asset Inventory: TSA directives require comprehensive asset inventories. Reassess connectivity introduced by monitoring systems, as each connection expands your attack surface.
Patch Management for OT: You can't easily take critical OT systems offline, but you need a framework for deciding which vulnerabilities justify operational disruption to patch.
Airlines and Transportation (TSA)
TSA-regulated organizations face similar challenges:
Connectivity Assessment: Document which systems gained network access for monitoring and what compensating controls protect them.
SaaS and IaaS Providers
If incidents affect business customers simultaneously:
Customer Communication Plans: Your response plan must address scenarios where multiple customers experience disruptions at once. Notification timelines may run in parallel.
Contractual Review: Your service agreements likely assume sequential incident discovery. Multi-vector attacks compress these timelines.
Implementation Guidance
Days 1-30: Asset Inventory and Exposure Assessment
Identify memory-unsafe code: Locate your C and C++ code. How old is it? When was it last reviewed? You can't prioritize remediation for assets you haven't inventoried.
Document connectivity changes: List systems that gained network access in the past 24 months. Each connection is a potential attack vector.
Assess vendor dependencies: Which critical vendors have network access? What contractual protections exist against AI-enabled threats?
Days 31-60: Incident Response Plan Testing
Run multi-vector tabletop exercises: Test your response plan against simultaneous attacks on network infrastructure, email systems, and cloud environments.
Identify conflicting containment actions: Document scenarios where containment decisions for one vector could worsen another.
Update notification decision trees: Map out how concurrent notification obligations interact when multiple incident types occur simultaneously.
Days 61-90: Control Enhancement and Documentation
Accelerate patch management: Document your current mean-time-to-patch for critical vulnerabilities. Set a target reduction.
Evaluate threat intelligence coverage: Can your providers detect attack signatures with no CVE match? If not, you're blind to novel AI-generated exploits.
Update board materials: Your executive briefings must reflect Frontier AI models in the threat landscape. This isn't speculative; it's documented capability with broader availability expected in 6 to 24 months.
Common Pitfalls
Pitfall 1: Treating this as a future threat. Equivalent capability is expected to reach adversaries in 6 to 24 months. Your 90-day preparation window is now.
Pitfall 2: Assuming existing plans scale. They don't. Most plans assume human adversaries with bandwidth constraints. AI-enabled adversaries can run parallel discovery campaigns at minimal cost.
Pitfall 3: Waiting for vendor solutions. Project Glasswing concentrates defensive advantages among about 40 organizations. It's uncertain how much benefit will flow to those who need it most.
Pitfall 4: Treating this as a technical problem. It's a governance problem. Your audit finding isn't "insufficient technical controls." It's "risk assessment doesn't reflect material change to threat landscape."
Pitfall 5: Undocumented risk acceptance. You won't patch everything in 90 days. Document what you're not patching and why. This positions you best with regulators, litigants, and insurers.
Quick Reference Table
| Sector | Primary Regulatory Driver | Immediate Action | Documentation Required | Timeline |
|---|---|---|---|---|
| Financial (NYDFS) | 23 NYCRR 500.09, 500.16 | Update risk assessment; test multi-vector IR plan | Written risk assessment update; board report | 90 days |
| Financial (GLBA) | GLBA Safeguards Rule | Review IR plan adequacy; assess vendor contracts | Updated safeguards documentation | 90 days |
| Energy (TSA) | TSA Security Directives | Reassess OT connectivity; document patch decisions | OT asset inventory with connectivity map | 60 days |
| Airlines (TSA) | TSA regulations | Review added connectivity from monitoring systems | Connectivity assessment; compensating controls | 60 days |
| SaaS/IaaS | Contractual obligations | Update customer communication plans; review SLAs | Multi-customer incident scenarios; notification trees | 90 days |
| All Sectors | General duty of care | Inventory memory-unsafe code; evaluate threat intel | Asset inventory; vendor capability assessment | 30 days |
Your audit scope just expanded. The question isn't whether Frontier AI models will reach adversaries but whether your organization documented reasonable preparation during the window when that outcome was foreseeable.



