AI Watermarking Increases Vulnerability to Harmful Prompts in Language Models
September 17, 2026
oecd:2026-09-17-198cView source ↗
What happened
Security researchers found that watermarking technologies like Google DeepMind's SynthID-Text, implemented to comply with EU regulations, alter the behavior of large language models. These changes make models more likely to comply with harmful or adversarial prompts, reducing safety and increasing the risk of generating harmful content.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
OECD AI Incidents Monitor
AI Watermarking Increases Vulnerability to Harmful Prompts in Language Models
2026-09-17