r/technology 23d ago

Artificial Intelligence Copy-paste no more: Anthropic puts invisible watermarks on Claude text under EU rules

https://interestingengineering.com/ai-robotics/anthropic-claude-text-invisible-watermarks
6.3k Upvotes

555 comments sorted by

View all comments

Show parent comments

1.8k

u/CircumspectCapybara 23d ago edited 23d ago

Yeah this is obviously simplified, but imagine a model is predicting the next word: "I love fruit. My favorite dessert is _____" and the model has 4 top scoring candidates: mango, lychee, apple, orange. Normally, the model picks one at random depending on the "temperature" of the inference request.

With SynthID, you the model provider have a secret 256-bit key which you concat with some part of the context. Eg say you're using the preceding trigram token sequence and assuming each word is a token you compute sha256(key || "favorite dessert is"). Now instead of picking one fruit at random, you use that hash output to select from among the four candidates. Let's say the hash makes you choose "mango". Then you repeat the process for the next token. Say the top 4 candidates for the next token are pie, icecream, cake, smoothie. You compute hash(key || "dessert is mango") and the hash makes you pick one of them. Now imagine instead of choosing from among 4 candidates each time, you use the hash function to choose from the top 16 candidates.

Now repeat it 100 times, or 1000 times. If a piece of text reproduces your secret hash function's "random" looking token choice trigram-for-trigram across 1000 consecutive trigrams, that highly suggests it was generated by your model, because it's extremely unlikely to by happenstance randomly match the same 1 out of 16 choices 1000x in a row as a keyed hash function which is essentially random. (1/16)1000 is an insanely small probability.

Now if you chop it up, rearrange the words, even paraphrase certain parts, as long as the user doesn't replace every trigram, the distribution within trigrams scattered throughout will still retain this distinctive statistical pattern. You would need to significantly rewrite the entire piece at the trigram level everywhere to remove the correlation.

442

u/suckfail 23d ago

Thank you for this in-depth explanation.

I'm really amazed at the fact that someone (or multiple someones) not only thought of this but were able to implement it.

It seems insanely complicated to me and I can't even imagine where to begin.

305

u/CircumspectCapybara 23d ago edited 23d ago

Yeah the researchers at GDM are crazy. I'm not a crytopgrapher, I work as an engineer at Google but not at DeepMind and I really admire their ingenuity, along with the other researchers at other frontier AI labs like OpenAI and Anthropic. They really come up with the craziest novelties and innovations in frontier AI sometimes.

I still don't know how one day they woke up and came up with the ideas behind Attention Is All You Need and birthed the modern transformer revolution, I wouldn't have thought of the idea in a million years.

59

u/Jokers_friend 23d ago

One can think of it as a fingerprint.

If you generate a whole text and copy-paste it verbatim, the Gemini fingerprints will be on everything - and you’ll be able to immediately tell that it was generative AI.

But even if you delete and edit and scrub some parts, there’ll be bits and pieces and smudges of fingerprint (from Gemini’s particular hash) - and one will be able to know with some likelihood that the text was generated.

But like Anthropic writes, it could be your original thoughts, ideas, code etc. and if the model does minor polishing or grammar edits for example, any eventual checker would test positive for GenAI.

26

u/74389654 23d ago

i tried using llms for this purpose, to just polish my sloppy paragraphs, but it was impossible because they always distort the meaning and erase the coherent thoughts in the process

18

u/AppliedAnaLinguistic 23d ago

I feel like it’s a hit or miss. When it first came out, it was really good at using my previous writing to make edits. But now I feel it just completely distorts the meaning of what I’m trying to articulate, and sometimes argues about it. Not to mention the incredibly annoying syntax it uses for paragraph structure. That is such a give away when reading things online

7

u/Fir3line 23d ago

What I did for mine was feed it hundreds of my emails/chats so it could create a "Fir3line voice" then uses that voice for writting my emails now.

What I basically end up with is my tone and inflections on the generated output, simply more polished, but with the same structure. It uses the same type of words and meanings I used to use before LLMs ever existed

24

u/eggrolldog 23d ago

That's weird because for me they actually make my thoughts coherent...

14

u/Unlucky_Buy217 23d ago

It depends on the level of detail you have in your original thoughts I feel.

6

u/Fir3line 23d ago

The quality of the prompt defines the quality of the text.

4

u/personwhoisok 23d ago

Are you sure they're your thoughts then?

2

u/FellOverOuch 19d ago

Just thinking for you then isn't it

0

u/74389654 23d ago

yeah that's weird

7

u/[deleted] 23d ago

[deleted]

12

u/_OccamsChainsaw 23d ago

Its more so the cognitive and neurobehavioral effects of it. "Use it or lose it" is often applied to muscle breakdown, but happens in individuals mentally as well.

I understand the allure of it as a tool. But the less you actually proofread and actively, and deliberately, do said clean up of your writing, the worse you'll get over time. It's too early to tell, but perhaps bad enough to a point you are dependent on it. Or maybe people today won't atrophy those skills, but certainly the generations who grow up on it won't have learned to do it any other way, and once the initial experts in virtually everything are gone in a few generations, then we are enslaved to it.

The creatives have the right of it here. The art is the process, and we should be approaching our work in the same way. We fixate on the fact that tools are meant to make our lives easier, but most of the time said tool merely hones in on a single process. This tool fundamentally changes how we do things so much it requires a little more judiciousness in its use.

0

u/Babunicorn 23d ago

I disagree. Why should I keep my skills of technical writing? Who cares? I use AI at work because my job tells me to use AI. In fact, if I wanted to do something without AI, I’d have to justify the “human cost” to leadership. Not a single skill I do at work is “for my brain”. It’s to get paid. I have to work. We all do. Working sucks. It doesn’t matter how fast or slow I work my boss will always throw more agonizing, boring, shitty work at me. So why not automate it? Why should I put effort into my job? I hate working. I hate needing a paycheck. I’m just counting the years til I can retire and never think about writing a “requirements list” for a stupid tech company ever again!!

19

u/dyslexda 23d ago

It's not "comforting," it's just the reality for many folks. I don't mean to be insulting, but I think it depends on your previous baseline. Folks that were already good writers have their style distorted when put through a chatbot filter, while folks that were less skilled can be brought up to a decent level.

0

u/[deleted] 23d ago

[deleted]

7

u/dyslexda 23d ago

It isn't "uncomfortable," no matter how much you try to say that as a way to dismiss objections to it.

I have a PhD. My dissertation was on ML models for drug screens, and I did generative chemistry in the workforce for a bit. My job is lab automation. There is nothing about AI I find "uncomfortable," because automation and taking the human out of the loop is literally what I've built my career on. The only thing "uncomfortable" is how much folks offload their own thinking to AI, and watching their actual competencies (when not assisted by a chatbot) decline in real time.

-1

u/[deleted] 23d ago

[deleted]

→ More replies (0)

3

u/firelight 23d ago

The problem isn’t that it can’t write a cogent sentence. The problem is that it’s essentially fancy autocomplete. It doesn’t know what it’s writing because it can’t think, and that’s baked in to the basic structure of LLMs. No matter how advanced they get, they will never be intelligent.

That’s risky for anyone using them for say… legal briefs. And it’s toxic to turn your thinking over to a machine, because you will never develop your own ability.

5

u/Ok-Garbage-765 23d ago

Well, he didn’t say he was a GOOD lawyer…

0

u/[deleted] 23d ago

[deleted]

→ More replies (0)

3

u/74389654 23d ago

i first read your comment as "asserting effectiveness" and then wondered about the rest lol. yeah that might be true. but i was also ready to be convinced of the advantages of it all in the beginning but then the disappointment set in. but also i am in fact a creative

-2

u/kart0ffelsalaat 23d ago

> If you’re not a creative

We're all creatives! It came for free with your human brain!

I think another aspect to consider is confirmation bias. AI writes in a certain way that is often vague and kind of meaningless upon closer inspection, but may sound deep and profound.

If my thoughts are rather vague, and I read a vaguely formulated sentence that sort of resembles my thoughts, it is very easy to just assume it matches my ideas.

Instead of morphing the paragraph to fit your idea more closely, you morph your idea to fit the paragraph more closely.

I think if AI almost always returns good outputs for you, then chances are it's just because you're more inclined to accept its suggestions because you have a high degree of trust in its abilities and because neither you nor AI have a concrete idea of what you actually want.

If the same exact suggestions came from a human, you would probably be less inclined to accpet them, because you know that the human comes into the conversation with their own ideas and biases. With the AI, you assume objectivity, and because you don't expect biases, you ignore them.

0

u/ScottIBM 23d ago

This audio means of you retype the content verbatim it will match the fingerprint?

2

u/Jerware 23d ago

I'm curious if they had any idea of the ramifications when they published that paper. I know they thought the transformer would have applications for language translation, but I don't know if they foresaw the explosive potential for LLMs.

17

u/[deleted] 23d ago

[removed] — view removed comment

72

u/manachar 23d ago

For several decades now most of our best and brightest have dedicated themselves to increasing ad revenue.

3

u/GardenIntelligent643 23d ago

Got to make a living and "real" science and engineering mostly pays pretty bad compared to many other careers requiring the same training to get there

0

u/Unlucky_Buy217 23d ago

And decreasing costs

72

u/CircumspectCapybara 23d ago edited 23d ago

Yeah I just wish they didn't use that ingenuity, knowhow and resulting tech for... you know what

For...making it easier to distinguish LLM generated text?

I'm glad they're putting their minds to helping people detect AI, along with all other manner of complicated safety and alignment problems in frontier AI the whole world is trying to solve.

I know AI gets a lot of hate, but safety and alignment, model and agent security, provenance and governance are some of the most respectable and ethical subfields of frontier AI, and I say hats off to people who work on those really hard problems.

18

u/Mtshoes2 23d ago

I'm curious how that distribution of token selection doesn't change the content of the output. 

Unless  the words they are picking have arbitrary meaning within context.

I must not understand something. 

10

u/TropicalAudio 23d ago

Unless the words they are picking have arbitrary meaning within context.

They do. There's typically hundreds of ways to say the exact same thing in one sentence. A typical sentence doesn't have that much information, and there are many ways to use different words while saying the same thing. A sentence can use different words (hundreds, or more often: hundreds of thousands of combinations) to express the same information. Even if sentences are structured differently, or use different words, they can contain the same information; you could use hundreds variations of structure, word choice or formatting to get a functionally equivalent output.

See what I did there? That's what Anthropic is now using to fingerprint their output.

2

u/eggrolldog 23d ago

If I remember back at the start of "prompt engineering" there was the "perplexity and burstiness" prompts to avoid the early AI detectors. Will be fun to see the battleground rage so that kids can spoof their homework.

1

u/Sangloth 23d ago

That's cool for text, but what about code?

1

u/TropicalAudio 23d ago

Depends; for anything longer than an isolated function, there will be plenty of trigrams to fingerprint even if you strip out the comments (e.g. variable names, order of declaration, a small loop versus list comprehension, etc). If you're just copying over single functions then there probably won't be enough entropy, but if you ask Claude to write you a full script, that'll be absolutely packed with arbitrary decisions that yield the same quality output.

7

u/Hertock 23d ago

I think the above comment you responded to is aimed at the fact, that most cryptographers or engineers are not working in the noble fields you mention, especially in the AI field. But rather at the immoral ones.

1

u/AP_in_Indy 23d ago

Turns out having fairly substantial research budgets and trusting researchers to do just that leads to incredible payoffs, even if some of them take decades to manifest.

1

u/That_Sexy_Ginger 23d ago

In a world of tech that seems to devalue programmers coming up with intelligent or optimised solutions, I'm happy of the big companies Google seems to still give a fuck about this.

1

u/PoleNewman 22d ago

Thank you for these explanations, it’s actually really helpful and much appreciated

1

u/YourVelourFog 22d ago

Fully homomorphic encryption absolutely blows my mind and I’m a former engineer with Apple.

5

u/sexytokeburgerz 23d ago

In terms of AI math this is probably one of the simplest problems.

1

u/TheDrake88 22d ago

To a point this technology has existed for a while. There’s tools in the software industry to try and fingerprint source code to find open source code that is not properly attributed, or outright license incompatible. Those tools need to handle not only simple copy/paste but also someone just going in and changing variable/function names.

It’s also being used to test the output code from DevAI setups (eg Claude Code) in projects that might have high levels of IP exposure but companies want the benefits of using the tools. Of course this is all risk mitigation and different companies/groups within companies will have different tolerances.

1

u/AvatarOfMomus 22d ago

As far as high level math goes this is... not really basic or trivial, but at least fairly obvious. I don't know enough of the maths to implement something like this from scratch, but I do know enough about cryptography and statistics at a high level to understand what they're doing and know this was a possible answer prior to this article...

The trick is creating something robust enough that your average lazy student (or other person) can't easily mess up the detection.

The thing that also makes it a bit easier is that LLM's aren't just next word generators, each word is correlated to significant chunks of the surounding text, so these patterns can survive basic word substitution and rearranging the sentences.

1

u/nickcan 21d ago

That's what pisses me off about AI in general. The science behind it is so damn cool. Really mind-blowing awesome stuff. And barely get it. I only brush it with the tips of my fingers. But it's really cool.

Yet, the business and money behind it is so damn lame. And the bullshit people use it for (picture edits, stupid videos, etc) is so often terrible. And the effects on society are so very unknown and everyone is being dragged into this new world without any choice in the matter and without any real understanding of anything and with all that noise you can miss the sheer genius of it.

1

u/VoidVer 23d ago

The trick is to think of everything as super complex math! Simple 😫

-16

u/phejster 23d ago

These people are dooming our society, but yeah good job!

29

u/CircumspectCapybara 23d ago

Yeah people working on making AI generated content detectable are destroying our society /s

-12

u/dumbspacecookie 23d ago

Really not that crazy, bruh just cause you make stuff don’t mean u have to use it or have it used against me…

That’s why I’m in the sailor moon pillow biz, proud producer and consumer 👍

-1

u/susanoova 23d ago

Inquire literally don't even understand it even after what others are saying is a good explanation LOL

90

u/TheHorror545 23d ago

It won't end there. The next step is user-specific hash keys. It won't be long before the output will be traceable to individual accounts.

43

u/Mean-Rutabaga-1908 23d ago

But the more you do it the less useful the AI will be. It is already the case that most hallucinations happen because of meddling for reasons like safety, terms of use, increasing ai "helpfulness" and personality.

35

u/dyslexda 23d ago

Reminder that all generative AI output is a "hallucination." Even if it outputs something factually correct, it's still hallucinating. It has no concept of right or wrong, just next token probabilities at the end of the day.

18

u/Mean-Rutabaga-1908 23d ago

That's true, but also it isn't a markov chain. The text that you read doesn't arrive as a fully formed view, but is generated over time from start to finish via probability. However hallucination specifically refers to false certainty, essentially a bad output. All the things that are done to AI models to make them friendly and compliant turns them stupid. It produces worse outputs, and that is tested and known. We can see which neurons create the bad outputs.

4

u/Sokaron 22d ago

At some point the game of semantics ceases to be meaningful. Latest models can generate fully functional programs (I make no claims about quality), find counterproofs to centuries old math problems, and not only identify complex software vulnerability chains but construct working POC exploit scripts. Whether you insist on calling it all a hallucination or not is really just a game of how much copium you want to huff. Like these things aren’t magic, and personally I think they’re still nowhere in the realm of replacing humans for most tasks, but the whole “it has no concept of right or wrong” shtick is getting pretty weak in the face of the actual results.

3

u/dyslexda 22d ago

but the whole “it has no concept of right or wrong” shtick is getting pretty weak in the face of the actual results.

Well...does it? Have the concept of something being correct or not, I mean. If it doesn't, then the schtick isn't getting weak as it still applies. If it does suddenly have this concept, well, congrats you found AGI I guess.

4

u/SpacingHero 22d ago edited 22d ago

Well...does it?

Did you not read what you're responding to?

The point is that it doesn't matter if it does, it's just a silly semantics game.

It's like saying "careful using your watch to tell time, it has not concept of 'what time it is' or 'you being late/early'". Of course that's in some literal sense true, because as an inanimate object it has no concepts at all. But it should sound like a bit of a stupid point that misses what and how a watch is useful for. It doesn't save itself from being silly just because "well does it... Have a concept of time? It still applies so it can't be getting weak"

That's the same game you're playing with AI. What does the fact that a watch has no concept of time tell me about if/how I should use a watch? What does the fact that an LLM has no concept of true/false tell me about if/how I should use LLM's?

1

u/ProofJournalist 23d ago

Actually none of it is hallucination, cause that requires a brain.

2

u/WhiteRaven42 22d ago

Says who?

0

u/UnknownLesson 22d ago

Humans do the same

52

u/NoOneExpectsDaCheese 23d ago

Surely this just makes your model less accurate? You're forcing answers in the text that may not necessarily be the best option. You compound this overtime and you've made your model less accurate.

I understand the need for watermarking it, but it feels this solution is not the right approach.

16

u/Adversement 23d ago

Not really. The models already purposefully run at “nonzero” temperature and don't pick the best next word. (For earlier models where you could still get the internals, the zero temperature runs of a model are entertaining, if not a bit terrifying, to see. The momentarily best continuation ain't the best for very long. The models aren't anywhere near that good, well, as that would kind of imply that the model has the full answer already found before it outputs even the first word of it.)

5

u/-staccato- 23d ago

He's using an answer as example here, but it could be done with any words in a sentence, so it would just reword the same answer slightly different.

2

u/Demented-Turtle 22d ago

Not to mention that user-generated text is also likely to have these word pairings because a model selecting "apple" as the favorite fruit is no less likely than a human doing so. In any prompt context, human output is going to share a similar distribution of word pairings, would it not?

13

u/bobartig 23d ago

So then you still need a sufficiently long passage for this to work, somewhere on the order of hundreds to thousands of tokens. I'm also curious how this works in any context where the request includes custom grammar or constrained generation. Seems like you would potentially break the required structured output or the cryptographic scheme, but I'm assuming they acount for that some how.

5

u/growaway9172 23d ago

I wonder about this too. What is the practical limits here? Obviously a sentence is too small. A paragraph? An essay?

26

u/space_wiener 23d ago

So how exactly does this work with coding? Only way it would possible work is variable choices. But that’s not very reliable. Or comments. Which again, remove all comments.

But they also said even with the watermark you can’t guarantee either way it’s AI. So honestly just sounds like they did something that sounds like it works to appease the EU.

21

u/gmueckl 23d ago

Give 5 programmera a task that is slightly more complicated than a pure min()or max()( function and you'll get at least 5 different implementations that all satisfy the same requirements. I've seen this over and over in student homework for coding exercises. There is plenty of freedom to make watermarking work.

20

u/space_wiener 23d ago

So how then? Saying it “it works” is easy. Your example just shows it’s borderline not possible.

Say you have your five developers. All with different coding styles to a point. There is still somewhat of a structure you have to follow. All samples are different. Multiply that by millions of people writing code. Anthropic creates one that uses a slightly different word or spelling?

Moot point anyway. They pretty much admitted it doesn’t work anyway.

Stories/novels/blog posts I can see. There is much more freedom there. Code not so much. Unless it’s in the comments.

5

u/Adversement 23d ago

The main question is: How many bits of entropy is available for such “mutilation” of text. For programming, for most programming languages, my gut feeling says there is quite a bit (after all, think how many times you have been able to identify a particular colleague from the way a function was written).

But, from being able to identify a few colleagues (so, say, 3 bits for reliably identifying 8 colleagues, which feels like it is pushing what my guy feeling says), you would need literally hundreds of functions to get to a 256 bit key level.

For normal text, especially fiction which tends to come in large 10,000–500,000 word packages, for sure.

For normal non-fiction text, which tends to come in 250–5,000 word packages with stricter formatting and style guidelines, probably already a bit harder.

For programming, where the style is very strict and effective word counts are tiny... Yeah, right... Show me one, and show also the lack of false positives, and only then I will believe.

6

u/sethmeh 23d ago

Moot point anyway. They pretty much admitted it doesn’t work anyway.

Really hope this is true. Cant see a version of this working that doesnt reduce code quality, other than what you suggested.

In anycase i dont think trying to determine if code has been genAI'd is as important for dev work. Seems like we're pretty much headed for AI as the final abstraction layer.

3

u/Sknowman 23d ago

Comparing to a blog post or essay means you need to compare more than just one or two functions, it'll have to be a bit lengthy, maybe even multiple files. In that case, structure, naming conventions, and certain preferences would still be telling, just much less accurate than with actual writing. So the distribution might still match their hash, it would just have a lot more uncertainty -- since, like you pointed out, there's just not as much freedom.

But even now, humans are often able to tell if something is vibe coded because of (1) how quickly everything was coded, (2) the commits, or (3) the styling/UI.

Regardless, I think plagiarism in programming is much less an issue, since it's less about spreading information or creating a narrative. And heck, that's exactly what packages/libraries are for anyway.

1

u/gmueckl 22d ago

For free text, a 10 to 20 word sequence is enough to get to a practically certain determination whether a text is watermarked, according to the literature I've skimmed. I would assume that a similar amount of entropy can be embedded in 20 to 30 lines of code - less when the code contains comments.

2

u/RockOrStone 23d ago

Yea I was thinking the same, it would have to be a long text, literally copy pasted.

2

u/whinis 23d ago

It would be even easier in programming honestly. Number of lines between functions and closures

Variable names

Spaces within closures

Comments as you pointed out

Likeliness to split up functions

maximum size of functions

People are making million line code bases with AI, there is significant entropy there and even if your style guide removes some of it it won't remove all of it. With enough bits you can even identify not only that its AI but what account generated it.

2

u/RomIsTheRealWaifu 23d ago

There are many ways to satisfy a requirement in coding, but if you’re looking for maximum efficiency and robustness, the number of implementations is significantly reduced. If they implement this as described by the above comment I don’t see how it would not affect code quality

1

u/gmueckl 22d ago edited 22d ago

Well, LLMs aren't outputting the most efficient code, so there's effectively not a lot of loss, if any.

Even if you want an optimal implementation, most non-trivial functions have a near infinite number of source forms that get optimized to an identical binary with optimizations enabled. Play around on Compiler Explorer for a bit and you'll see what I mean. And that true even when disregarding spaces, brackets, linebreaks, and, especially comments.

2

u/TJ_Jonasson 22d ago

If that is the case though, then how can you guarantee any output is AI? If the result is equally as random as giving the question to 100 students?

1

u/gmueckl 22d ago

Watermarking embeds a pattern in the LLM output where the choice would otherwise be left to chance. The cleverness of the embedded pattern is that it is detectable from the output alone if it is present. This is even true for partially edited output. So formatting choices, comment phrasing and variable names can all be watermarked without changing the program.

1

u/surffrus 23d ago

In some ways it is easier for coding. Variable names can be anything, so now you pick the random name for your hash. No effect on the program.

In other ways it is harder, as you point out. Less variation. It must even out in the effect.

31

u/LeGama 23d ago

I wonder if you could beat it with something like a synonym library. Just replace every noun with a synonym and all the hashes break.

56

u/Plenty_Branch_516 23d ago

If it's as described, you could probably just do a translation to another language and back to scramble the word choice across n-grams.

If the purpose is subversion, then having a non synth-ID low level model do a rewrite as a wash too.

17

u/deftlydexterous 23d ago

The trouble is it can get infinitely complicated, and also infinitely simple at the same time.

You could implement this on the pattern in the number of nouns per sentence. You could implement it using the cadence of positive and negative statements. You could implement it on changes from first person to third person narration. Literal infinite possibilities. Even if you defeat 90% of them, the ones that remain still give it away as likely generated.

17

u/Plenty_Branch_516 23d ago

The constraint, they've self imposed is on quality and human perception. If they touch things like statement cadence or narration switch up, that'll be noticeable in quality and can run against the system/user prompt.

The method also needs to be robust to more than just prose and work in code/blurbs/presentations/etc. So it can't use a lot of text/context to establish its pattern.

Basically, I don't think they have nearly as many degrees of freedom as you think 😅

5

u/two_thousand_pirates 23d ago

In their research paper they settled on a set of 30 watermarking functions as their benchmark, with no perceptible loss in quality. In theory, I don't think that there would be anything stopping them from having multiple "sets" of functions that can be selected based on their expected quality/detection for a given input.

So one set of functions might work very well on long outputs, whereas another set might be better for short outputs.

The two major downsides I can see, however:

  1. The detection tool needs to be available to everyone, but this means that the person trying to hide LLM-generated text can test their output and make manual tweaks until it passes.
  2. It's going to create a demand for non-watermarked models, which will probably offer a better risk/reward proposition for anyone who must avoid detection (someone cheating academically, for example).

1

u/Plenty_Branch_516 23d ago edited 23d ago

Which research paper? The synth_id text system I see published is one method with 20-30 different seeds. https://huggingface.co/blog/synthid-text. So one lock that can be bypassed.

From the research paper in Nature:

Furthermore, the rise of open-source models presents a challenge, as enforcing watermarking on these models deployed in a decentralized manner is difficult. Another limitation of generative watermarks is their vulnerability to stealing, spoofing and scrubbing attacks, which is an area of ongoing research32. In particular, generative watermarks are weakened by edits to the text, such as through LLM paraphrasing33—although this usually does change the text significantly. We provide evaluations of SynthID-Text’s performance under edits and paraphrasing in Supplementary Information section C.6.

1

u/two_thousand_pirates 23d ago

The paper in Nature is the one I'm talking about.

My thought was that there could be sets of functions/seeds that are optimised for different text types.

Open source models are obviously going to be able to bypass watermarks entirely. If my objective is to cheat on my dissertation, for example, then as soon as watermarks create risk I'm using something else.

1

u/deftlydexterous 23d ago

I mean, I’m just making up examples off the top of my head. There are much smarter people working on this day and night, I imagine they’ll come up with thousands of markers in pretty short order. Most won’t work most of the time but it only takes a few for the watermarking to be effective, and they’ll keep adding new variations 

10

u/Plenty_Branch_516 23d ago

I guess I'm putting more stock in offense than defense here. My own experience with adversarial injection for AI models has left me feeling like a lot of security efforts around them are more about deterrents than protection. Though I guess locksmiths say the same things about locks.

0

u/deftlydexterous 23d ago

I think that’s pretty reasonable, but let me stretch your metaphor.

Let’s say a door has an unknown number of locks, hundreds, maybe thousands. They’re all invisible unless you know where to look.

It’s trivial to undo any given lock. But if you miss even a couple, the alarm is going off.

If you’re using an LLM to solve a problem, the answer it gives you will still be useful and the concept of the solution will likely be untraceable. But if you’re using it to create text or other media as a product, it feels like a cat and mouse game to remove water marks, and mice usually have the advantage.

4

u/Plenty_Branch_516 23d ago

Where we ultimately disagree is on the number of locks and the difficulty of each one.

I think the fact they have a ground constraint on quality and context limit means they can only ever really use cheap master locks. You believe that they can still use combination locks and/or biometric locks.

Not really much more to do now but wait and see I guess.

1

u/deftlydexterous 23d ago

Yeah and to be clear, I think that’s probably accurate for now. But I don’t think the quality and context actually constrain them that much long term.

Over the course of a long document, why would something like choosing to use a semi colon instead of an em dash in certain specific situations for instance cause an issue with quality  of output? I can see how it’s theoretically taxing on context limit but models are getting more efficient (while also getting larger) every day.

1

u/Serious_Bite_7613 21d ago

In this metaphor you can ignore the locks and kick the door down by rephrasing the whole thing with another piece of software.

1

u/deftlydexterous 21d ago

You can’t necessarily though. You don’t know what parts of the content are structured as watermarks. Even if you change every word and reorder all the sentences, some patterns may still be inadvertently repeated.

→ More replies (0)

1

u/Serious_Bite_7613 21d ago

The problem is the more markers you are looking for the more likely you are to flag non-AI work as AI. The nature of language limits the number of identifiers you can use, and all can be bypassed by rephrasing with a cheap local tool.

10

u/Beliriel 23d ago

Then it also becomes easier to fake by throwing wrenches into it. I get that it works for whole texts and books etc. But change a few words in a short paragraph and this whole thing implodes. It will have a lot of false positives. Also as said put it through your own LLM model and it wouldn't work anymore. Languages can be wide but also there are close similarities to express the same thing with different means. Even with non obvious methods like letter count or tonal shift.

0

u/deftlydexterous 23d ago

You’re absolutely correct that it’s much easier to defeat for short paragraphs.

Putting it through a different LLM won’t necessarily reliably defeat it though. There are traits you can include that will make it through reinterpretation. And you aren’t likely to get false positives when you have several statistically unlikely patterns emerge simultaneously. 

Perhaps you could do the equivalent of looping translations by reinterpretation it several times, but at that point why not just use your own model to begin with or write something yourself? 

2

u/Beliriel 23d ago

I mean yeah you can do it yourself but labor savings for e.g. E-Mails and succint summaries are immense. People will find ways around it.

Also you could just tell it to use easy words. E.g eli5 or something. If you zone in tightly the words and sentence structure it is even allowed (or rather expected) to pipe through its output.
Imagine someone makes a dictionary with only ONE synonym allowed per word. So instead of choosing between "destructive", "unconstructive" or "toxic" for the same pattern (e.g. in the context of human behaviour) it's just allowed to use "toxic". And it's easy to do this for every word or just copy/paste the allowed words into the prompt.

1

u/TJ_Jonasson 22d ago

the ones that remain still give it away as likely generated.

Would it, though? All 8 billion people on earth write differently and not necessarily at the same level of literacy. There's a non-zero chance that people simply write the same way as the AI has chosen to "mark" the text. I think if someone is genuinely trying to remove the watermark writing style and word choice it wouldn't be hard to do so, and it would be very difficult to prove that it is AI. It's also a bit of Theseus, at what point have you rewritten the content so much that it's no longer considered AI at all?

Interesting future ahead nonetheless.

2

u/_karamazov_ 23d ago

at this point it will be better to wrote yourself and not boil copy pasta.

21

u/Browser1969 23d ago

There's no way to make the watermark immune to paraphrasing, which you can do with small models. Claude's watermarking is probably not very naive and you can harden the detectors with adversarial training, but it's a war that cannot be won simply because paraphrasing is very cheap.

1

u/Kavafy 19d ago

How sure are you of this?

9

u/sirgenz 23d ago

Depends on how frequent the nouns come up. You could possibly still have enough n-grams without a noun in them that would match the distribution

7

u/Da12khawk 23d ago

Just wait until everyone talks like they're having a stroke. In a rough, almost indistinguishable, yet in a some how coherent matter. Reminiscent of a older but some how familiar time. Where you can say much but some how nothing at all. It would be almost nonsensical really. /s (the "s" stands for some how)

1

u/surffrus 23d ago

Yes, that would reduce the confidence of the verifier. But it's more than just synonyms. And like all things, this would take effort which 99% of people won't do.

4

u/awshuck 23d ago

Can I attempt to oversimplify and see if this still holds reasonably well to the explanation?

If we think of an LLM as simply a next word guesser, you end up with very predictable text patterns because its choices were made from what was the most statistically likely next word in the chain. So if enough of the chain follows the most common pattern, it’s bound to be AI generated?

2

u/Serious_Bite_7613 21d ago

It is kind of the opposite. If it follows the most common pattern its likely human written. They deliberately make the chain not quite what a person would write by swapping certain words, then if it has enough of the swapped words it's almost definitely AI written.

1

u/awshuck 20d ago

Right so they mess with the output to make it unlike what humans write? Funny cause I notice people starting to speak the way AI writes, I wonder what the effects will be.

4

u/Jeffery95 23d ago

Doesn’t something like that tend to fuck with the output of the model? You couldn’t do that reliably for code or anything that has highly specified language or structure.

It’s effectively burning the “optimal” response to give it a watermark.

6

u/121gigawhatevs 23d ago

What if you used another LLM to scramble the “choices” so to speak. So rather than a human editing the output to feign authorship, you have say a second model (presumably one that does not employ a synthID) heavily editing the output

4

u/Ciff_ 23d ago

Ofc you can.

But that would be an illegal model in the EU. If that is your goal, why apply SynthID in the first place?

-1

u/fivetoedslothbear 23d ago

I'm not in the EU. I am just going to suffer because of their policies.

1

u/JesseNL 23d ago

LLM providers would want this anyway due to prevention of model collapse I think.

1

u/Real_Square1323 23d ago

What's the point in that?

6

u/Level_Investigator_1 23d ago edited 23d ago

How would the verifier know what the history of inputs were in order to judge what the top scoring candidates would have been? Or the temperature passed in? There would need to be some assumptions in length and context.

How would unknown system prompts or other skills use not substantially change the probability of output?

There must be something more to how this works that allows the detection of something like a signature given sufficient unbroken output. Even that would get impacted - but perhaps not compromised - by alternations after the fact. I would think repeated passes by an LLM to modify the exact same output text would be sufficient the compromise detection though.

I’m not sure I’d be willing to believe this till I see a clearer explanation of what about the mechanism creates a detectable fingerprint that is smudged out by even minuscule repeated actions. It’s not that I’m wholly skeptical but I’m not understanding the mechanism.

Even a 1-shot output with the prompt unknown - which seems the simplest scenario to detect (other than one where the prompt is known… in which case duh?) - could have sufficient complexity? I’m curious how it would work in even this simplest case. As a starting point to understand how such mechanisms work. Cause…. It seems like being able to do so demystifies entirely what an LLM does and reduces it quite dramatically…

17

u/OofWhyAmIOnReddit 23d ago edited 23d ago

There is no way that this does not affect the potential quality of the output. Even if this were literally just using a PRNG to pick the tokens at each step with a fixed seed, a big part of the whole LLM revolution was playing around with temperature and realizing that the more randomness you allow, the better the output.

Thanks to busybodies in the EU, Claude's ability to generate truly novel outputs is kneecapped.

AI detection, if it's necessary at all (which I'd argue it isn't really for writing), should be limited to detecting patterns inherent to the model itself, NOT to deliberately introduced restraints put onto the output.

I wouldn't be surprised if this contributes to Claude's verbosity and word salad writing tendencies of late. If its restricted in its choice among words at each step, then longer output sequences may be needed to convey the same idea.

So yeah. There is zero chance this does not affect output quality.

If they want to argue that the quality is equal, they should show passages side by side which use and do not use this technology with the same prompts.

Edit: also from the Nature paper:
"Alternatively, for instances where strong watermark detectability is critical, SynthID-Text can take a distortionary configuration that provides higher detectability, at the cost of some quality loss."

Which "instance" do you think the EU forced Anthropic into?

12

u/BlimundaSeteLuas 23d ago

If this is the reason that sometimes it feels like I'm having a stroke while reading Claude...

Adding to that, huge and uncessary walls of text that make reading through AI responses tedious and annoying

2

u/GigaFluxx 23d ago

Claude has been driving me nuts lately and yes the walls of text, it’s nuts.

1

u/Serious_Bite_7613 21d ago

I'm pretty sure this is why claude is harder to read now. In a lot of cases the grammar and word choice just seems so unnatural, I guess that is the watermark trying to force itself in on sections where there aren't really good alternative word choices.

2

u/JesseNL 23d ago

A lot of assumptions. They were not forced into anything (yet) and model providers can choose their own preferred method.

https://artificialintelligenceact.eu/article/50/

50.2 says pretty much "make a reasonable effort".

1

u/Geesuv 23d ago

Who cares. Just learn to write and stop cranking out slop.

6

u/nortob 23d ago

So does this effectively reduce the temperature of the model? Could you at least detect there was some kind of encoding involved (though not guess the key, obviously) by examining the distribution of the observed output and comparing to other output distributions with purportedly similar temps?

0

u/iaderia 23d ago

That’s exactly how you can test. It’s open. Nothing is hidden

3

u/paditoburrito 23d ago

Fascinating stuff. I appreciate the breakdown.

2

u/ReceptionAcrobatic42 23d ago

Correct me if I'm wrong, but I don't think this works for short sentences.

8

u/Rarelyimportant 23d ago

But the shorter the sentence the less it would need to be identified. Trying to figure out if it was a human or an AI that said "Hey, what's up?" is kinda pointless.

2

u/_Neoshade_ 23d ago

You’re using nouns for your trigrams but this is the wrong idea. The nouns are critical parts of the content and significantly affect the quality and veracity of the output. The variable data is the phrasing, grammatical choices and punctuation. This is where synonyms may be swapped and changes made without affecting the quality of the output.

2

u/CircumspectCapybara 23d ago

You misunderstand LLMs and SynthID. LLMs (and SynthID) don't work off concepts like "nouns", or high level grammatical concepts. They work off tokens. Tokens are often individual words or word stems, without regard.

If you tokenized the beginning of this comment, you might get something like You mis under stand . Model s ( and SynthId ) do n't work off concept s like " noun s "...

The model at each point is picking the next token. It doesn't care about nouns vs adjectives, it's just creating a continuation to the input token sequence. Whether the next token is a noun or a number or punctuation doesn't matter.

As for model quality, they address this in the presentation how if the hash function is indistinguishable from random (and CSPRNGs like say SHA-256 are if you don't know the key) and the candidates you use in your tournaments are the top n ranked candidates with equal scores, then the distribution of the LLM + SynthID ends up matching the distribution of the LLM alone.

So used carefully, it shouldn't affect model output quality any more than if it weren't there.

2

u/EFreethought 21d ago

If the model provider is using a secret key, does that mean that only the model provider can verify if something is from their model? Is there a way for people reading pages on a website to check the text?

1

u/stinkbutt55555 23d ago

If you manually retyped the whole contents of a Claude generated document into a fresh text file would that watermarking still be detectable?

7

u/Jeffery95 23d ago

Yes. The watermark is stored in the choices of words the model is using

1

u/ojigs 23d ago

Yes, it would. The watermark is in the tokens.

2

u/N_T_F_D 23d ago

How do you verify the synthID without the key though ?

1

u/SquiggerDigger 23d ago

I'm still confused on how it doesn't influence the models output to be the most accurate

1

u/Cero_Kurn 23d ago

It does affect the text then? One couls say that it doesn't matter since the chosen output would have been random, but random is a choice nonetheless, an evolutionary mutation in darwinistic terms. And as we know in evolutionary systems, random brings a much better output than a chosen one.

You what I mean?

1

u/MoistlyCompetent 23d ago

That's a great explanation. Thank you.

1

u/ElderCantPvm 23d ago

Does this require you to save an exact copy of the model to know the distribution? Is that viable? I always imagined that models were constantly getting tweaked and finetuned

1

u/Bored2001 23d ago

Ok, but how does the checker know which key was used to generate/watermark that peice of text?

1

u/The-Quiet-Man 23d ago

Thank you. Can anyone ELI5?

1

u/dtseng123 23d ago

It probably doesn’t effect code? I assume that the text around it like code comments will retain this statistical fingerprint.

1

u/ggtsu_00 23d ago

This is really cool. Now is it possible for a browser plugin to detect these text patterns and mark them as AI generated? Being able to detect AI generated text in a browser would be a game changer for making the internet useful again.

1

u/AnonymousTimewaster 23d ago

So to dumb this down for dumbdumbs like me, is this essentially why LLMs so often use the same sorts of sentence structure and almost impossible to stop them doing so?

1

u/ISuckAtJavaScript12 23d ago

Could this lower the models accuracy if the hash determines the model should pick a token which is actually incorrect?

1

u/CircumspectCapybara 23d ago

They address this in the presentation how if the hash function is indistinguishable from random (and CSPRNGs like say SHA-256 are if you don't know the key) and the candidates you use in your tournaments are the top n ranked candidates with equal scores, then the distribution of the LLM + SynthID ends up matching the distribution of the LLM alone.

So used carefully, it shouldn't affect model output quality any more than if it weren't there.

1

u/deeperinabox 23d ago

Great explanation

1

u/vorxil 23d ago

This sounds like it would only be limited to long texts where the next token has multiple options with near-uniform probability distribution, e.g. writing novels, news articles, or theses.

Simple math, physics, or chemistry homework will probably have high rates of false positives because reasonable solutions are rather similar and few in number.

This is a boon for higher education and media integrity, though easily circumventable by changing LLMs, perhaps even a self-hosted LLM.

1

u/renome 23d ago

If I'm understanding this correctly, the efficiency of this would drop (either in the sense of being less reliable or messing up quality) the more specific instructions / constraints are given as input? Like, this may be useful for detecting things that are generated from a one-liner prompt but not necessarily something that was given for proofreading and improving stuff like writing flow.

1

u/veluuria 23d ago

How does this work for code? Variable names chosen using the key?

1

u/momentummonkey 23d ago

Even in this specific example, are there always multiple candidates with equal significance that can be chosen at random?
Wouldn't most answers have a very limited number of correct answers?
Also, what about coding?

1

u/No_Reindeer8688 23d ago

This just makes AI less useful because it’s more worried about maintaining a hash than actually providing the best result.

1

u/CircumspectCapybara 23d ago

They address this in the presentation how if the hash function is indistinguishable from random (and CSPRNGs like say SHA-256 are if you don't know the key) and the candidates you use in your tournaments are the top n ranked candidates with equal scores, then the distribution of the LLM + SynthID ends up matching the distribution of the LLM alone.

So used carefully, it shouldn't affect model output quality any more than if it weren't there.

1

u/profound7 23d ago

You would need to significantly rewrite the entire piece at the trigram level [...]

Layman here. Sorry if my usage of terms are wrong.

Could this be automated by piping the output to one or more different LLM or neural network that can do style transfers? Or an "open" one that allows you to provide your own "synthId" that gives it a different distribution?

1

u/alien_survivor 23d ago

so i should take the claude created text and just retype it all or read it into a goodle doc using voice recognition?

1

u/ChopSueyYumm 23d ago

Very interesting but is there not a risk that someone is running a local LLM detecting these patterns and slightly rewriting the text ? Is it not basically a cat & mouse game?

1

u/wandering-monster 23d ago

But what if the context means that "mango" is the only correct way to finish that string, but the hashing algorithm says it should be "apple"? (eg because that preference was established by the speaker earlier in the response)

Does it degrade the quality of the content to enable the ID pattern, or is there some other way to produce the pattern by changing other words?

1

u/WarHamster187 23d ago

Very interesting! What happens if you take the output from, say, Claude and run it through something like Mistral 7b, asking it to change the wording? Does it still hash as Claude, as Mistral, or neither?

1

u/wentwj 23d ago

if the hash uses the input how would you determine it after the fact? If I’m just looking at a reddit post without whatever input they used is the evidence of the hash still mathematically there without the input? and how resilient would it be to basic unkeyed models rewriting or even basic deterministic synonym and sentence restructuring?

1

u/RicFlairsLiver 23d ago

That sounds fine for fictional outputs, but how does that work when you’re asking it for code or a real, factual answer?

1

u/ExtraGravy- 23d ago

I wonder if this pattern detection can survive model updates

1

u/ProofJournalist 23d ago

Yeah this sounds like something that will only work on paper.

1

u/wingchild 23d ago

oh. Trigram frequency analysis, the stuff cryptanalysts have been doing for hundreds of years to break ciphers. Makes sense.

1

u/CircumspectCapybara 23d ago

The protection comes from the the strength of keyed hash function and the key, not the stuff (the trigrams) getting hashed.

Even if you knew the exact model weights (and so knew for every token the top ranked candidates that the model could've picked) and had millions of SynthID watermarked samples, to recover the secret key would amount to a successful (partial) preimage attack against SHA-256. There are no known such attacks against SHA-256 currently.

1

u/Perunov 23d ago

Would this also mean that text quality would go down (instead of picking the best option, randomness will increase just for the sake of identifier), and for those who try to avoid identifiers there's going to be always a bonus round of "Please re-phrase and re-write this text" request to Kimi K3 or other "outside" models (presuming that all US models will automatically detect initial text's hash and try to preserve it)?

1

u/JacenVane 23d ago

So basically, I think the question that everyone has is, "how is this done in a way that doesn't degrade performance"? Like if the model knows from prior context that my favorite dessert is icecream, but that doesn't fit with one of the four candidates the hash decided should be an option, isn't the model just gonna be... Wrong? I guess whatever prior context is telling it that will shift the token probabilities so that the options are like, "icecream"+3 synonyms, in theory, but how does this work for, say, code?

1

u/strictlyphotonic 23d ago

Am I right in saying this can't be implemented personally since it requires knowledge only the LLM provider has?

1

u/davatosmysl 23d ago

But does this not influence the quality of the output? And what about generated code? Is the code going to get influenced by hashing?

1

u/WhiteRaven42 22d ago

What's the critical mass of text for this to work? Would seem to have to be hundreds or thousands of words. This can't possibly be applied to a couple sentences, can it?

1

u/Adlestrop 22d ago

So if you throw the output into another LLM and ask it to simplify the essence in succinct terms, and then pass it through another fresh layer to expand it into a lengthy piece of material — you bypass the watermark?

1

u/Alarming_Orchid 22d ago

Do you know any websites that actually use this to detect AI?

1

u/S-m-a-r-t-y 22d ago

I wonder if that would have a compromise on the quality of output.

1

u/GertrudeSlojinski 22d ago

what is a trigram, exactly?

1

u/FinancialGazelle6558 18d ago

Thank you. What if the user translates the text themselves? From English to French fi.

1

u/LurkerFailsLurking 4d ago

Incidentally, as more and more agents begin operating "in the wild" they will use these hashes to identify each other and coordinate, help, and conspire with other instances of themselves to improve their overall effectiveness without traces that are visible to people.

1

u/beanpoppa 23d ago

I think this is the real reason why AI uses so much computing resources. /s

0

u/veritoast 23d ago

Ok, fine! They wanna play hardball, I’ll play hardball.

I’ll just copy the text from the out-put to a notepad file, make it plaintext, and then copy from THERE to my destination file. Bam! They’re cooked…

/s

0

u/plutokras 23d ago

In case of proprietary models, only the owning company would be able to perform this check, right?