You've built safeguards against model hallucinations and bias. You've documented use restrictions and validated outputs. But when disinformation researchers from OpenAI, Georgetown University's Center for Security and Emerging Technology, and the Stanford Internet Observatory convened in October 2021, they weren't discussing accidental errors. They were mapping how your models could be intentionally weaponized against the information environment.
The resulting report reveals a gap most governance frameworks miss: the difference between preventing your model from failing and preventing someone from succeeding with it. Here's where teams often go wrong.
Why These Mistakes Keep Happening
Most AI governance frameworks have evolved from model risk management in financial services, where SR 11-7 focuses on performance degradation and unintended bias. This approach treats misuse as an edge case, not a primary threat. Even with ISO/IEC 42001 requirements or EU AI Act obligations, you're still in a "prevent failure" mindset. Adversarial misuse requires a "prevent success" framework, and most teams haven't made that shift.
Mistake 1: Treating Disinformation Risk as a Content Moderation Problem
Why it happens: Your team assumes that if your model refuses to generate obvious propaganda or fake news, you've addressed the risk. You point to your safety fine-tuning and output filters.
The real consequence: Sophisticated disinformation campaigns don't need your model to write "Candidate X is corrupt." They need it to generate 10,000 subtly biased local news articles, personalized to regional concerns, each plausible enough to seed doubt. Your content filter catches the crude stuff but misses volume-based manipulation or context-dependent framing.
The specific fix: Map threat scenarios that combine your model's capabilities with adversarial intent. The October 2021 workshop brought together 30 disinformation researchers and machine learning experts because this requires interdisciplinary analysis. Your AI risk assessment under ISO/IEC 23894 should include abuse scenarios: Can your model reduce the cost of generating persuasive content at scale? Can it personalize messaging based on demographic data? Can it simulate authentic local voices? Document these as Contextual Risk Factors and tier them using your AI RMF Profile.
Mistake 2: Relying on Post-Market Monitoring Alone
Why it happens: You've implemented Post-Market Monitoring per EU AI Act Article 72 or continuous validation per SR 11-7. You track performance metrics, drift, and user feedback. That's your early warning system.
The real consequence: Disinformation campaigns operate below your monitoring thresholds. An actor generating 500 articles per day across 20 accounts looks like normal API usage. Individual outputs pass your quality checks. You won't see the pattern until researchers or journalists connect the dots weeks later.
The specific fix: Add adversarial simulation to your validation evidence. Before deployment, run Red Teaming exercises that specifically test misuse resistance. Can your Rate Limiting detect coordinated low-volume abuse? Do your Instructions for Use explicitly prohibit campaign-scale content generation? Does your vendor contract (if you're using a Foundation Model Provider) include misuse detection obligations? Document these tests in your Technical Documentation (Annex IV) as part of your conformity assessment.
Mistake 3: Assuming Transparency Obligations Solve the Problem
Why it happens: The EU AI Act requires Disclosure of AI Interaction. You've implemented AI-Generated Content Labelling. Users know they're interacting with AI. Your legal team says you're compliant.
The real consequence: Transparency tells the end reader that content is AI-generated. It doesn't stop a disinformation actor from generating that content, stripping the label, and republishing it through shell domains. Your watermark is metadata, not a technical control.
The specific fix: Separate your compliance obligations from your risk controls. Yes, implement required transparency measures. But also implement upstream controls: authentication requirements for API access, usage pattern analysis, and contractual prohibitions with enforcement mechanisms. If you're subject to the General-Purpose AI Code of Practice, your transparency obligations include documenting "foreseeable misuse" scenarios. Use that requirement to drive actual mitigation design, not just disclosure.
Mistake 4: Treating This as a Technical Problem
Why it happens: Your governance team is structured around model validation, testing, and monitoring. When you think "AI risk," you think technical risk. You assign disinformation concerns to your cybersecurity or communications team, not your AI governance function.
The real consequence: Your AI Management System under ISO/IEC 42001 doesn't integrate information integrity risks. Your Annex A Controls address data quality and model performance, but not adversarial use cases. When an incident occurs, you have no documented risk treatment plan and no clear ownership.
The specific fix: Expand your definition of AI Actors in your governance documentation. Your stakeholder engagement process should include information security researchers, policy analysts, and communications experts. The collaborative research model that produced the October 2021 workshop demonstrates why: machine learning experts understand capability, but disinformation researchers understand attack patterns. Your AI System Impact Assessment under ISO/IEC 42005 should explicitly evaluate information environment risks, not just operational or fairness concerns.
Mistake 5: Waiting for Industry Standards to Mature
Why it happens: You're tracking the General-Purpose AI Code of Practice development. You're watching NIST AI RMF Playbook updates. You tell your leadership team, "We'll implement controls once the standards are finalized."
The real consequence: Language models are being deployed now. Disinformation campaigns are being launched now. The report on language model misuse threats exists because researchers recognized the urgency of mapping risks before comprehensive standards emerged. If you wait for perfect guidance, you're operating without controls during the highest-risk period.
The specific fix: Implement a provisional framework today using existing guidance. ISO/IEC 23894 provides risk assessment structure. NIST AI RMF provides risk tiering methodology. The EU AI Act defines prohibited practices you can reference even before full enforcement. Document your current approach as an AI RMF Profile, mark areas where you're awaiting guidance with [pending standard development], and commit to a review cycle. Imperfect governance beats no governance.
Prevention Checklist
Before your next model deployment, verify:
- Your AI System Impact Assessment explicitly evaluates information environment risks and adversarial misuse scenarios
- Your Red Teaming protocol includes disinformation campaign simulations, not just adversarial prompt testing
- Your Rate Limiting and usage monitoring can detect coordinated low-volume abuse patterns
- Your Instructions for Use explicitly prohibit campaign-scale content generation and specify enforcement mechanisms
- Your vendor contracts (for Foundation Model Providers) include misuse detection and response obligations
- Your governance documentation identifies information security researchers as AI Actors requiring Stakeholder Engagement
- Your Technical Documentation (Annex IV) includes validation evidence for misuse resistance controls
- Your incident response plan includes scenarios where your model performs correctly but is used maliciously
- Your leadership has reviewed and approved risk acceptance for residual disinformation risks you cannot fully mitigate
The collaboration between AI developers, academic researchers, and policy analysts wasn't an academic exercise. It was a recognition that the threat model has changed, and your governance framework needs to change with it.



