Back Issues/Search Home → Calendar → Archive → RSS → Subscribe → Current Issue → Popular →

All issues › Volume 342, Issue 5 › IT News › AI

How LLM Watermarking Can Change AI Agent Behavior

TechTalks, Monday, September 28th, 2026

Lasso Security research finds SynthID-Text watermarking can shift agents' tool choices, arguments and refusals.

Lasso Security tested Google DeepMind's SynthID-Text watermarking across several open models and found it can change which tools agents select, the arguments they pass and whether they refuse harmful requests, an effect it calls sampling drift.

Watermarking lowered tool-call accuracy on six of seven models, but the more telling metric was churn: on average 6.5% of individual task verdicts flipped, and up to 16.8% for phi-4, even when headline accuracy barely moved, with effects growing under prompt injection.

The researchers urge providers to publish tool-calling and safety evaluations when watermarking changes, and teams to re-test agents when deployment settings change.

more →  ·  More from AI →