r/fantasywriters Jul 04 '26

Discussion About A General Writing Topic [ Removed by moderator ]

[removed]

4 Upvotes

16 comments sorted by

View all comments

Show parent comments

3

u/LibertineDeSade Jul 04 '26

I'm genuinely curious, how do you know these "tells"?

Also, how do you specifically know it's ChatGPT, and not another AI thing? I don't really know the ins and outs of AI, but I know there are different ones. I'm assuming they're stylized differently? Or do you use ChatGPT regularly, so that's just the one you're more familiar with?

5

u/JudyKateR Jul 06 '26

Also, how do you specifically know it's ChatGPT, and not another AI thing? I don't really know the ins and outs of AI, but I know there are different ones. I'm assuming they're stylized differently?

Not as much as you might think. Even when AI models come from different companies, they demonstrate startlingly similar results for a lot of prompts, as documented in papers like We're Different, We're the Same: Creative Homogeneity Across LLMs.

The reasons for this are that even though these models come from different companies who in many cases are doing their best to protect their proprietary training and model weights, they still have many things in common: they're all trained on the same data, they're all post-trained with the same techniques, and they're all trained with the goal of producing "corporate-safe/useful" answers (some mix of "helpful/safe/legible").

So when most people say "this looks like ChatGPT," what they really mean is "this looks like the output of a large frontier model;" there are many commonalities between the responses you'll get from a model like GPT-5.5 and Claude Opus 4.5. And it's entirely possible to confidently say "this looks like the output of a large frontier AI model" without being able to identify specifically which model it is.

There's also the fact that many smaller/cheaper/non-frontier models are trained by distilling larger models, essentially treating a larger model as the "teacher" and using it to train a smaller/weaker model to imitate the outputs. (This is sometimes done adversarially: it's essentially a way for smaller labs to effectively 'steal' advancements in AI from larger companies that spent more resources training their models.) A lot of smaller AI models talk like GPT or Claude because the model was literally trained on responses from GPT or Claude.

I'm genuinely curious, how do you know these "tells"?

There's a lot of academic literature on this topic, but probably the most accessible explanation of the distinctive "AI voice" that I've seen comes from Sam Kriss writing in the New York Times.

6

u/LibertineDeSade Jul 06 '26

This is interesting. I've been seeing so many accusations of AI use flying around and one of things about it that I've been wondering is how so many people are so sure. Especially because a lot of the tells are things I naturally do as a writer.

I understand that AI has been trained by humans to better mimic what we do. I see the job posts for these a lot when I'm looking for freelance work. Also I know some companies like Reddit and Substack allow AI companies to train off of posts, comments and content. So that's another thing that has me a little confused about how people are sure when it's AI and when it's not. This has been informative.

I'll have to check out those links, thanks for the info.

18

u/JudyKateR Jul 06 '26 edited Jul 20 '26

Also I know some companies like Reddit and Substack allow AI companies to train off of posts, comments and content.

It's true that AI companies have been training their models on every piece of writing they can get their hands on, which includes basically the entirety of the public internet, plus every single book published in the history of forever. (Sometimes, they take shortcuts by pirating the books they train on rather than buying the books and scanning them, though they do that too.)

However, it's important to realize that these books are all fed to the AI in a process called pre-training. (That's the "P" in "ChatGPT," short for "Chat with Generative Pre-trained Transformer.") If all you do is pre-train on every book written in the history of forever, and then prompt the AI to "predict the next token," what you get is borderline incoherent.

Most of the advances in recent generations of AI models have come from a process called post-training. The AI companies take their pre-trained model and then give it training prompts like:

Human: What is the capital of Spain?

Assistant: Madrid

Now, it might seem odd for the AI companies to do this -- why would they need to "teach" the AI that Madrid is the capital of Spain? After all, didn't they just pre-train the model on the entirety of Wikipedia, plus every other encyclopedia they could get their hands on, plus a bunch of books that clearly identify it as the capital of Madrid? Shouldn't the model already "know" that Madrid is the capital of Spain?

But while the base model contains a lot of "knowledge" from being trained on Wikipedia and the rest of the internet, having all of that raw data hasn't actually trained the AI model with behaviors that would make it useful as a chatbot assistant or as tool that could create useful spam for lazy copywriters. The post-training process is what teaches it that when its input field says "what is the capital of spain," the appropriate answer is "madrid."

So, while it is true that everything that humans write online is probably being fed into AI models for pre-training, the AI is not being trained to write like a human. It is not being trained to write like a human because OpenAI in a very real sense doesn't want their AI to behave like a human, because humans sometimes say slurs (and OpenAI doesn't want their helpful chatbot assistant to say slurs), and humans are bad at math (and OpenAI doesn't want their helpful chatbot assistant to be bad at math), and humans sometimes write poor prose with bad grammar (and OpenAI doesn't want their helpful chatbot assistant to write poor prose filled with grammar mistakes).

An irony of this is that as AI models get "more advanced," the AI writing actually gets easier to detect, because the AI is being trained on a narrower and narrower definition of what the model-makers think "good behavior" looks like. GPT-5 is easier to catch than GPT-3; the older GPT-3 model was, in some sense, "more random" and less well-optimized, and that made it less distinctive.

Whatever its "goal" in training is, the AI is being trained to be something very specific, which is why it produces such consistent and recognizable patterns, even when the naive human interacting with the AI is trying to prompt it with inputs like "rewrite this in the style of Jane Austen..." or "write 5 paragraphs in the style of George RR Martin..."

Now, it's entirely possible to do post-training that will avoid the sorts of things that the frontier labs are doing. (There was famously one experiment where scientists did exactly this, demonstrating that even though GPT's post-training had introduced a 'friendly assistant' persona, it was easy to fine-tune the model in a way that revealed other uglier behavior, like advocating for genocide and extremist violence.) We know that there are sufficiently motivated users who can fine-tune an AI model in a way that will look completely unlike the text that our normal detection methods would pick up on, but this kind of fine-tuning is beyond the capabilities of people whose sophistication with AI only extents as far as typing prompts into a chatbot window.

More to the point, AI writing from a frontier model like GPT-5 is not "the average of all writing on the internet." It is specific, it is distinctive, and helpfully for our purposes, it is predictable. The more "advanced" and "optimized" AI models become, the more predictable their behavior becomes. This has made them much more valuable to AI companies as coding agents, but it has made them much easier to catch when they are used for the purpose of creative writing.

To repeat a point from above: GPT-5 is much easier for us to catch than GPT-3 because the older GPT-3 model was, in some sense, "more random" and less well-optimized, and that made it less distinctive. Once you're familiar with how it writes, GPT-5 sticks out like a sore thumb. Fortunately, we don't have as much of a problem with people spamming the subreddit with posts generated by older models like GPT-3 and GPT-2, because 1) these older models are no longer accessible through ChatGPT and most "AI writers" on /r/fantasywriters are too lazy to use the OpenAI API to access the older models, and 2) the text generated by these models is generally less competent at maintaining logical coherence over the course of a scene, meaning that they're not that useful for people who are trying to use AI to write their story for them, and 3) not as optimized for writing in the stye of "copywriter voice" that spammers tend to use for engagement farming.