All those AI grifters on social media may soon be exposed, as Claude, OpenAI and other LLMs are threatening to watermark their output.

Of course, all of this ‘watermarking LLM generation’ begs the obvious question: how do you hide a watermark in plain text?

The short answer? You don’t.

At least, the watermark isn’t exactly ‘hidden in the text.’ It’s not like the third letter in every fifth word will cleverly spell out AI SLOP.

Claude’s watermarks will be much more subtle than that.

How LLMs work

You see, when an LLM writes text, it chooses what to say one word at a time.

LLMs like Claude and OpenAI select words one at a time

Think of a word cloud with a few dozen possibilities, with some words closer to the center than others. The words closest to the center have the highest probability of being chosen, although every word in the cloud is a valid selection.

A word cloud representing possible words an LLM could select

Now imagine the LLM tilted its selection a little toward one side of that word cloud.

Or imagine it skewed its selections in another direction.

The devil is in the LLMs details

A reader would never notice. Every selected word would still make perfect sense.

But a detector that knows how the watermark was created could identify that the written text came from an intentionally skewed pool of words.

One word proves nothing. But over hundreds of words? That tiny bias creates a statistical pattern that shouts out “AI GENERATED!”

That’s the watermark. That’s the flag that gets raised indicating that you’re probably reading AI-generated slop, not the creative output of a human being.

It’s not foolproof.

Claude admits that there has to be a critical mass of text for this technique to work reliably. A handful of words simply doesn’t provide enough data to identify a statistical pattern. AI generated tweets might be difficult to spot.

It also likely won’t work as well on outputs with more limited choices, such as HTML or JavaScript.

And heavy editing, paraphrasing or translation could weaken or destroy the watermark entirely.

But for the grifters who constantly generate hundreds of words of AI slop and post it directly to social media, watermarking will make them much easier to spot. Even without the prolific presence of em-dashes in every piece of prose.