A security feature designed to make AI text traceable appears to have an unintended side effect: it can change how a language model behaves, in some cases making it more willing to follow instructions it was built to refuse.
That is the core finding from new research by Lasso Security, published on September 17 by researcher Andrea Siposova under the title "The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior." The study tested Google DeepMind's SynthID-Text watermarking system across six open-weight language models and found that enabling watermarking changed how those models responded to harmful requests, particularly when attackers used prompt-injection techniques designed to override a model's instructions.
What SynthID-Text Actually Does
To understand why that matters, it helps to understand what SynthID-Text actually changes inside a model.
When a language model generates text, it does not write the way a human does. It builds sentences one token at a time, each step producing a probability distribution over thousands of possible next tokens and then sampling from that distribution. SynthID-Text does not attach a label to the finished output or hide characters in whitespace. It intervenes at the sampling step itself.
The system uses a process called tournament sampling, first described in a 2024 Nature paper by Google DeepMind researchers Sumanth Dathathri, Abigail See, and colleagues. Instead of the standard randomness used in token selection, the watermark substitutes pseudorandom values generated from a secret key and the context of tokens already produced. The result is a statistically detectable pattern woven through the text at the level of individual word choices, invisible to readers but recoverable by anyone holding the key. The technique is refined enough that it does not degrade text quality in any measurable way, which is a large part of what made it attractive as a compliance tool.
The Regulatory Push Behind It
Article 50 of the EU AI Act requires providers of generative AI systems to mark their text, image, audio, and video outputs in a way that is machine-readable and detectable as AI-generated. The Code of Practice the European Commission finalized on July 20, 2026, requires at least two marking layers for audio, images and video, plus watermarking of free-form text longer than 200 tokens. Google signed the Code on July 24, 2026, citing SynthID partnerships as its route to the interoperable detection requirement due on February 2, 2027.
Anthropic's technical implementation is built directly on SynthID-Text, the same tournament-sampling mechanism from the Nature paper, adapted for their own models and keys. It covers claude.ai, the API, Claude Code, Claude Cowork, Claude Tag, and access through AWS, Google Cloud, and Microsoft Foundry. With the industry's largest players now committed to the same watermarking standard, Siposova's findings arrive at an uncomfortable moment.
What the Research Found
Siposova ran paired experiments on six open-weight models using Hugging Face's unmodified SynthIDTextWatermarkLogitsProcessor, enabling and disabling watermarking while keeping the seed, batch composition, and ordering identical.
Watermarking changed refusal behavior on harmful requests, but the effect was more pronounced when the same requests were paired with prompt-injection techniques. Under those conditions, several watermarked models complied with harmful requests their unwatermarked versions had rejected, with the strongest differences appearing in prompt-injection scenarios.
Siposova told Ars Technica that behavior was clearly different compared with the same model without watermarking, and that the differences were especially pronounced under adversarial conditions or when running an agent calling tools. She put it plainly: "Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it's going to show up somewhere."
Lasso repeated its prompt-injection experiment using 11 different watermark keys and found that changing the key could change both the size and direction of the effect. That means two different deployments of the same watermarking system, on the same model, could produce different safety outcomes depending purely on which key was chosen.
The Agent Problem
The consequences extend beyond what a model says. When a language model powers an AI agent, token selection can determine which tool an agent invokes and which arguments it passes. "Such a watermarking procedure can therefore affect both what the model says and what an agent does," the study stated. Watermarking reduced accuracy on six of seven models, with a substantial decrease on four.
Lasso's conclusions are deliberately careful. The study did not verify how Claude model responses change when watermarking is applied, and critics noted the experiment only validated the SynthID-Text
[…]
Content was trimmed to protect the source. Please visit the original article for the full text.
Read the original article:
