A 56% increase in AI-driven attacks. That figure comes from IBM’s Cost of a Data Breach Report 2026, and it’s the number I keep coming back to whenever someone asks me why anyone bothers with AI safety research in the first place. People are already pointing these systems at targets, and the systems are already being poked at in ways their builders never planned for.
So when a claim started circulating that large language models respond differently to harmful prompts when AI watermarking is in play, my ears perked up. It’s a juicy idea. It also happens to be one I can’t verify, and I want to walk you through why that distinction matters more than the headline does.
What’s actually being claimed
The short version of the story goes something like this. Watermarking is a technique for tagging AI-generated output so it can later be identified as machine-made. The claim is that when watermarking is switched on, a model’s behavior shifts when someone asks it for something dangerous — maybe it refuses more often, maybe less, maybe it answers differently.
Here’s where I have to be straight with you. I went looking for the specifics, and they aren’t there. The sources I can point to don’t document how LLMs respond to harmful prompts under watermarking. No measurement, no study, no numbers. The claim is in the air, but the evidence hasn’t landed.
That’s not me dismissing it. Plenty of real research starts as a hunch that turns out to be correct. But on a site about explaining AI agents to people who don’t build them, I’d rather tell you “we don’t know yet” than dress up a guess as a finding.
What we do know about how these models work
Large language models like ChatGPT, Google Gemini, and Anthropic Claude are trained on enormous amounts of text to produce writing that reads like a person wrote it. They’re evaluated on things like alignment (does the model do what we actually want), safety (does it avoid causing harm), and fairness. One of the main tools for testing them is red-teaming, where humans deliberately try to break the model to find its weak spots before someone with worse intentions does.
The other thing to understand is that these models are probabilistic. They aren’t looking up answers in a database. They’re predicting what text should come next, token by token. That’s why they sometimes produce confident-sounding answers that are flatly wrong — what the field calls hallucinations.
That probabilistic nature is exactly why a claim about watermarking changing harmful-prompt behavior is plausible enough to take seriously. If output is generated through a process of weighted choices, and watermarking works by nudging those choices in detectable patterns, then yes, it’s reasonable to wonder whether the nudging has side effects on safety behavior. Reasonable to wonder. Not the same as demonstrated.
Why this matters for anyone using AI agents
If you’re building or using an AI agent — something that takes actions on your behalf rather than just chatting — you’re stacking behaviors on top of each other. The model’s safety training, whatever filtering sits around it, whatever logging or tagging your vendor applies. Each layer interacts with the others, and those interactions are where surprises live.
The practical takeaway isn’t “turn off watermarking” or “panic about watermarking.” It’s that safety features are not independent switches. Changing one part of how a model generates text can ripple into parts you weren’t thinking about. That’s an argument for testing your own setup rather than trusting that each feature does only what its name suggests.
How to read claims like this one
A few questions I ask myself, and that you can borrow:
- Is there a named study, a dataset, or a measurement behind this, or just a confident assertion?
- Who tested it, and on which models? Behavior that shows up in one model often doesn’t transfer to another.
- Does the claim specify what “differently” means? More refusals? Different wording? Those are very different outcomes.
- Is the person making the claim selling something that the claim conveniently supports?
AI moves fast enough that unverified claims travel further than verified ones. The watermarking-and-harmful-prompts question is a good example of a story that could turn out to be important, and right now sits in the “interesting hypothesis” pile.
If solid research lands on this, I’ll cover it properly, numbers and all. Until then, the honest answer is the useful one. Somebody noticed something worth looking into, and the looking hasn’t been done in public yet. With 56% more AI-driven attacks on the board, I’d like that work to happen sooner rather than later.
🕒 Published: