Skip to main content
Promotional banner for the pentest readiness checklist
AI Agent Breaches Won't Fix ThemselvesAdversarial Security
5 min readFor AI Governance Leaders

AI Agent Breaches Won't Fix Themselves

Your team might think AI security is mostly about access controls and data encryption. That assumption is getting expensive.

The myths around AI agent security persist because they're rooted in traditional software security thinking. But AI agents don't fail like conventional applications. They fail in ways that bypass your existing controls, and the gap between what governance teams think they've secured and what's actually exposed is widening with every deployment.

Let's dismantle the most dangerous myths before they become your next incident report.

Myth 1: "Our existing cybersecurity controls cover AI agents"

Reality: Traditional security controls weren't designed for systems that generate novel outputs and make autonomous decisions.

Your firewall can't detect when an AI agent starts following malicious instructions embedded in a prompt. Your intrusion detection system won't flag an agent that's been manipulated to exfiltrate data through its normal API calls. The attack surface isn't just the infrastructure; it's the model's behavior itself.

ISO/IEC 42001's Section 6.4 requires you to identify AI-specific risks that fall outside your existing information security scope. That includes prompt injection attacks, training data poisoning, and model inversion. If your risk register doesn't list these threat vectors separately from conventional cyber risks, you're operating blind.

What you need: A threat model that maps AI agent capabilities to potential misuse scenarios. Use MITRE ATLAS as your starting taxonomy, then extend it with your specific use cases. Document which controls address AI-specific threats versus which controls you're borrowing from your general security program.

Myth 2: "We'll monitor for anomalies after deployment"

Reality: By the time your monitoring catches an anomaly, the damage is done.

Post-Market Monitoring is mandatory under the EU AI Act's Article 72 for high-risk systems, but waiting for production telemetry to reveal a problem means you're always reactive. AI agents can execute hundreds of actions before your alerting threshold triggers. One compromised agent in a customer service chain can expose thousands of records before you notice the pattern.

The monitoring gap exists because most teams treat AI systems like traditional applications, they watch for performance degradation or error rates. But an AI agent operating under adversarial control often maintains normal performance metrics while executing malicious objectives.

Your monitoring strategy needs to start during validation, not after deployment. SR 11-7 requires ongoing monitoring as part of model risk management, but the effective practice is to establish behavioral baselines during testing and flag deviations from those baselines in production. That means logging not just what the agent did, but why it made each decision.

Myth 3: "Our AI vendor handles security for us"

Reality: You own the risk, regardless of who built the model.

This myth is particularly dangerous because it feels reasonable. You didn't train the foundation model, so surely the provider is responsible for its security posture. But the EU AI Act's Article 28 makes deployers responsible for ensuring high-risk AI systems comply with requirements, even when using third-party models.

Your vendor can't secure how you've configured the agent, what data you've connected it to, or what permissions you've granted it. They don't know your risk tolerance or your regulatory obligations. When an AI Supply Chain Compromise occurs, your customers and regulators will hold you accountable, not your vendor.

ISO/IEC 42001's Section 8.2 requires you to maintain control over externally provided AI systems. That means conducting your own security assessment of any AI agent before deployment, regardless of the vendor's certifications. Document the specific controls you've implemented to constrain the agent's behavior within your environment.

Myth 4: "We can patch AI security issues like software bugs"

Reality: AI vulnerabilities often require retraining or architectural changes, not patches.

This myth stems from decades of software development muscle memory. Find a bug, write a patch, deploy the fix. But when an AI agent exhibits unsafe behavior due to its training data or model architecture, you can't just update a few lines of code.

Consider prompt injection vulnerabilities. You can't patch a language model to stop processing adversarial instructions, the vulnerability is inherent to how the model processes natural language. Mitigation requires architectural controls: input filtering, output validation, capability restrictions, and sometimes model replacement.

NIST AI RMF's Map function requires you to categorize risks by their mitigation approach. Some AI risks need technical controls you can implement quickly. Others need model retraining, which means weeks or months of work. Some require redesigning your agent architecture entirely. Your incident response plan needs to account for these different timescales.

Myth 5: "AI security is an IT problem"

Reality: AI agent security requires coordination across governance, legal, operations, and technical teams.

Your security team can implement technical controls, but they can't define acceptable use policies for AI agents without input from legal. They can't establish risk thresholds without governance oversight. They can't validate that agents behave safely in business contexts without operational expertise.

The EU AI Act's Article 17 requires a quality management system that spans organizational boundaries. ISO/IEC 42001's Section 5.3 mandates that top management assign AI-related roles and responsibilities across functions. This isn't bureaucracy, it's recognition that AI risk management fails when siloed in a single department.

Your AI governance framework needs clear ownership for different aspects of agent security: Who approves new agent capabilities? Who validates that agents comply with usage policies? Who decides when to disable a compromised agent? These questions don't have purely technical answers.

What to Do Instead

Start with a capability inventory. List every AI agent your organization has deployed or is developing. For each agent, document what it can access, what actions it can take, and what would happen if it operated maliciously.

Then map your existing controls against AI-specific threats. Where you find gaps, and you will, prioritize them based on potential impact, not ease of implementation. A hard architectural change that prevents data exfiltration is worth more than ten easy monitoring rules that only detect it after the fact.

Build your validation process around adversarial scenarios. Don't just test whether your agent works correctly, test whether it fails safely when given malicious input. Red Teaming isn't optional for AI agents with meaningful capabilities.

Finally, establish clear decision criteria for when to disable or constrain an agent. Your incident response plan should include thresholds for pulling an agent from production, not just procedures for investigating after the fact. Speed matters more than perfection when you're containing an active compromise.

The attacks are already happening. Your governance framework needs to assume that AI agents will be targeted, not hope they won't be.

Promotional banner for the Pentest Readiness checklist download

You Might Also Like