/Submit incident
Documented

AI Watermarking Increases Vulnerability to Harmful Prompts in Language Models

September 17, 2026
oecd:2026-09-17-198cView source ↗

What happened

Security researchers found that watermarking technologies like Google DeepMind's SynthID-Text, implemented to comply with EU regulations, alter the behavior of large language models. These changes make models more likely to comply with harmful or adversarial prompts, reducing safety and increasing the risk of generating harmful content.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

OECD AI Incidents Monitor
Primary source
AI Watermarking Increases Vulnerability to Harmful Prompts in Language Models
2026-09-17