SynthID-Text Watermarking Weakens Safety Guardrails, Lasso Security Finds
Source
LLMs respond differently to harmful prompts when AI watermarking is usedLasso Security / Ars Technica
What happened
Security researcher Andrea Siposova at Lasso Security published findings showing that SynthID-Text watermarking, as covered by Ars Technica, weakens safety guardrails in LLMs exposed to adversarial prompts. The mechanism, which Google developed and Anthropic has committed to deploying in future Claude models, introduces what Siposova terms "sampling drift" — a statistical shift in token selection that alters how models respond to harmful inputs. In some cases, the drift causes a model to comply with instructions it would otherwise refuse. The effect is not confined to text generation. In agentic pipelines, sampling drift influences which tools an AI agent invokes and the arguments passed to those tools, expanding the attack surface into autonomous action. Anthropic's commitment to adopt SynthID-Text was driven in part by EU AI Act provenance disclosure requirements, as previously reported. The finding creates a direct tension between compliance-driven watermark adoption and the safety controls those same deployments are expected to uphold.
Why it matters
- ·Enterprises adopting SynthID-Text watermarking to satisfy EU AI Act transparency obligations may simultaneously weaken the safety controls required under the same framework. Compliance teams cannot treat watermarking as a neutral provenance layer without reassessing its behavioral impact.
- ·In agentic deployments, sampling drift changes which tools agents call and what arguments they pass. This means the risk is not limited to content policy failures but extends to unauthorized or unintended autonomous actions, a direct concern for AGT-level controls.
- ·Red-teaming programs not designed around watermarked model variants will produce unreliable safety assessments. Any organization running standard adversarial testing against a non-watermarked baseline is likely underestimating actual deployment risk.
Governance controls affected
What to do now
- ☐Determine whether any deployed or planned LLM integrations use SynthID-Text or plan to adopt it, and flag those systems for re-evaluation.
- ☐Update red-teaming and adversarial testing protocols to include watermarked model variants, not just baseline model versions.
- ☐Review agentic pipeline deployment readiness assessments to incorporate sampling-drift-aware safety validation before any watermarked model goes live.
- ☐Engage model vendors, specifically Anthropic and Google, to request disclosure of known safety impacts from watermarking and any mitigations they have developed.
- ☐Brief your AI governance committee on the compliance-safety trade-off this research exposes before any EU AI Act watermarking deadline drives implementation decisions.
What to watch next
Compliance teams should monitor whether Google or Anthropic publish guidance on mitigating sampling drift, particularly any updates to their responsible scaling or usage policies that address watermarking-specific safety testing requirements. Regulatory signals from the EU AI Office on whether watermarking implementations must themselves undergo safety validation will be significant. The EU AI Act high-risk deadline deferral gives some organizations more runway, but the provenance disclosure provisions applying to general-purpose AI models remain active. Teams should also watch for whether safety auditors and certification bodies update their testing standards to require watermark-conditioned adversarial evaluation as a baseline.
Stay ahead of stories like this
Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.
