Anthropic recently added watermarks to its AI outputs in response to the EU AI Act transparency rules. Governance teams quickly began treating watermarks as the definitive test for AI-generated content. This mirrors the downfall of perimeter security in cybersecurity, and it's about to happen again in AI governance.
The issue isn't with watermarking itself. It's the mistaken belief that one trust signal can replace continuous verification. If your team is building AI content policies around watermarks alone, you're repeating the same mistakes that kept network security teams chasing breaches for decades.
Myth 1: Watermarks Tell You What's AI-Generated
The Reality: Watermarks indicate what the creator chose to disclose.
A watermark is a voluntary flag. It relies on the model provider implementing it, the user not removing it, and the content remaining unchanged after generation. That's three potential failure points before you even start checking for accuracy.
Consider when someone uses an AI tool to draft a section, edits it heavily, then publishes it. Is that AI-generated? Does the watermark survive the editing? These aren't edge cases. They're how people actually work with AI tools.
The EU AI Act requires transparency for General-Purpose AI Models, but knowing the tool used doesn't guarantee the output's trustworthiness. You wouldn't accept a calculator's output without checking the formula. Why accept AI content without verifying the claims?
Myth 2: Missing Watermarks Mean Human-Written Content
The Reality: Absence of evidence isn't evidence of absence.
Not every AI tool uses watermarks. Not every user keeps them. As John Kindervag's zero trust principle reminds us: never trust, always verify. The absence of a watermark doesn't prove human authorship any more than a badge proves someone works in a building.
This matters for compliance teams evaluating content for regulatory submissions, model documentation under SR 11-7, or Technical Documentation (Annex IV) requirements. You can't build an audit trail on "we didn't see a watermark, so we assumed it was human-written." That's not a control. That's a hope.
If your verification process stops at checking for watermarks, you're not verifying. You're just looking for a flag that may or may not be there.
Myth 3: Watermarking Solves the AI Disclosure Problem
The Reality: Disclosure and verification are different challenges.
The EU AI Act's transparency requirements push providers to disclose AI generation. That's useful for awareness but doesn't solve the verification problem, because disclosure is voluntary and easily circumvented.
Watermarks can be stripped, cloned, or applied to human-written content to create false positives. Anyone determined to misrepresent content origin will find a workaround, just as attackers found ways around perimeter firewalls.
Treating watermarks as the solution is perimeter thinking. It assumes one credential, checked once, establishes trust. Network security moved past this model because it failed. AI governance shouldn't repeat the cycle.
Myth 4: You Can Build Policy Around a Single Trust Factor
The Reality: Single-factor trust models don't scale under adversarial pressure.
Badge cloning, social engineering, and tailgating proved that one credential checked once isn't worth much. Security teams learned to verify continuously, regardless of who you are or where you came from. Content verification needs the same approach.
Instead of asking "does this have a watermark?" ask the questions that matter for any content: Is it accurate? Does it cite verifiable sources? Does it hold up when you check the claims yourself? These questions work whether a human or a machine did the writing.
For teams building AI Management Systems under ISO/IEC 42001, this means designing verification workflows that don't depend on watermarks as the primary control. Watermarks can be one data point among many, but they can't be the single test of trustworthiness.
Myth 5: Penalizing AI Use Will Stop Misuse
The Reality: Prohibition drives workarounds, not compliance.
People use AI tools because they increase productivity. Penalizing AI use doesn't stop the behavior. It just pushes it underground, where you can't govern it.
This is the lesson from zero trust cybersecurity: the bad actors find the workaround. You can't build security by banning tools. You build it by verifying outputs, regardless of the tool used.
Your governance framework should assume AI tools are in use across your organization. Design verification processes that work whether content comes from a human with a keyboard or a human using an AI assistant. The tool doesn't determine trustworthiness. The verification process does.
What to Do Instead
Stop treating watermarks as a trust boundary. Start treating them as one signal in a multi-layered verification system.
Build your AI content governance around these principles:
Verify claims, not credentials. Check factual assertions against source material. Require citations for regulatory submissions and model documentation. Don't accept content at face value because it lacks a watermark.
Design for continuous verification. One-time checks don't catch errors introduced during editing or integration. Build review cycles into your workflow, especially for high-stakes content like Technical Documentation (Annex IV) or validation reports under SR 11-7.
Assume tools are in use. Your teams are using AI assistants. Design policies that govern the output quality, not the tool choice. Focus on accuracy, completeness, and auditability.
Layer your controls. Watermarks can flag AI generation. Human review catches logical errors. Citation checking verifies factual claims. Audit trails document the verification process. No single control is sufficient.
The zero trust lesson is simple: verify everything, all the time. Content deserves the same rigor. If you're building AI governance policies around watermarks alone, you're not ready for what's coming.



