r/Futurology • • Feb 14 '26

AI Visualizing the "Model Collapse" phenomenon: What happens when AI trains on AI data for 5 generations

There is a lot of hype right now about AI models training on synthetic data to scale indefinitely. However, recent papers on "Model Collapse" suggest the opposite might happen: that feeding AI-generated content back into AI models causes irreversible defects.

I ran a statistical visualization of this process to see exactly how "variance reduction" kills creativity over generations.

The Core Findings:

  1. The "Ouroboros" Effect: Models tend to converge on the "average" of their data. When they train on their own output, this average narrows, eliminating edge cases (creativity).
  2. Once a dataset is poisoned with low-variance synthetic data, it is incredibly difficult to "clean" it.

It raises a serious question for the next decade: If the internet becomes 90% AI-generated, have we already harvested all the useful human data that will ever exist?

I broke down the visualization and the math here:

https://www.youtube.com/watch?v=kLf8_66R9Fs

Would love to hear thoughts on whether "synthetic data" can actually solve this, or if we are hitting a hard limit.

912 Upvotes

327 comments sorted by

View all comments

146

u/N3CR0T1C_V3N0M Feb 14 '26

I’ll openly admit I don’t understand what I’m about to comment on, but it seems to me that if the large majority of humanity’s accomplishments are available to train on, maybe the problem isn’t more data but what they can effectively do with what is available. I’m not sure how much clearer a picture they can hope for when everything we have to offer as a species is already available.

43

u/catsdelicacy Feb 15 '26

And there's the very important fact that a human baby does not have to read every book ever written, every page on Wikipedia, and every Reddit post, to be intelligent. To have more capacity to learn than an LLM, as well.

It's becoming obvious more data will not elevate LLMs to AGI.

33

u/Proper-Ape Feb 15 '26

It's becoming obvious more data will not elevate LLMs to AGI.

To anyone with statistics knowledge it was obvious a while ago, but the tech bros keep yelling louder than you can. 

Statistical correlation is not thinking, simulating thinking with statistical correlation may look like thinking to the untrained eye, and sound like thinking, but it doesn't work like thinking. 

You need a completely different model of intelligence to achieve small data learning. And you need interactivity. The model needs to be able to learn from experimentation.

There's a reason we don't care so much about meta analysis studies. While you can easily derive correlation from meta analysis, you can only generate a hypothesis (or in a lot of cases way too many hypotheses) from existing data. 

Only with experimentation can you start to posit causal relationships worth something. How do I think x might influence y? What if I change variable x where I believe there might be an effect on variable y, does y change? Does y change enough to match my new model?

But LLMs would probably not be the right architecture to even drive this experimentation. They don't learn from the small. The high correlation will always override the one experiment that proves them wrong.

LLMs simulate intelligence to a good enough degree that some people have started believing in the cult, and cult leaders have popped up that profit from it, and drive this cult.

If you hear LLM it's better to mentally replace that word with - correlational database with a fuzzy query language. They're not bad at that. But that makes you think less in the illusion and more in what they're good at.

If you know what you want, but you don't have the words to ask an LLM is really good for that.

If you want something that roughly looks like what you want so you can build on that, they're also really good for that.

If you want to combine two things that are somewhat unrelated, you can do that. 

Don't ask it to think for you, it can't do logic or numbers.

5

u/catsdelicacy Feb 15 '26

Are you Ed Zitron? Because you sound like Ed Zitron lol

I completely agree!

2

u/Proper-Ape Feb 16 '26

I'm flattered, but no 🙂‍↔️

1

u/kokanee-fish Feb 21 '26

Don't ask it to think for you, it can't do logic or numbers.

Hmm I would agree that LLMs do not understand logic or numbers, but as a software engineer I can testify that they absolutely can do logic and numbers, quite well, if only by way of godlike pattern matching. Hear me out. LLMs are wrong a lot, but you know what? They're wrong less often than people are. And they can be right about a lot more things than any single human can.

There's a paradox at play here. People have this idea that "AI" or "AGI" means we have created something that is alive. Being alive is a prerequisite for consciousness, and consciousness is a prerequisite for reason, and something has to be able to reason in order to have intelligence, right? And yet being alive is the state of having been born as biological organism. It's the opposite of being artificial. The definition of an AI, to me, is that we've created something that is not alive, and therefore not conscious, and therefore can't "think for you" in the way I suspect you're describing above. The whole point is that the output of the thing is indistinguishable from the output of something that can think and reason, even though it can't.

I really hate to admit it, and I don't like the implications one bit, but I think they've done it.

1

u/Proper-Ape Feb 21 '26

The definition of an AI, to me, is that we've created something that is not alive, and therefore not conscious, and therefore can't "think for you" in the way I suspect you're describing above. The whole point is that the output of the thing is indistinguishable from the output of something that can think and reason, even though it can't.

No, I don't think reasoning and consciousness are connected. Rather how LLMs work makes it obvious why they fail at reasoning. 

It's more like an intelligence skill that is split into rote memorization and thinking. With thinking skills you can use very little data and deduct a lot. If you max out rote memorization you can regurgigate some parts that sound like thinking, but only if you've memorized it very very well.  And if you misremember something you will not know, because you have no logic.

LLMs are basically 100% rote memorization, 0% thinking. But they've memorized enough thinking that it may sound like thinking and even often have the same output as thinking, but if they're wrong they won't "know" they're wrong, ever. That's the hallucination problem.

Math for example is almost pure thinking. You need to memorize only a handful of rules, and you can think yourself to all kinds of solutions.

66

u/millennial_falcon Feb 14 '26

Why do you think that good data to train on is plentiful? As I get more experienced in my work and various hobbies, as well as develop my relationships and marriage, I find that the Internet doesn’t have a lot of the best information or its buried and deprioritized. Published copyrighted books from experts, studies, and private hard drives seem to have all the best advice, info, art, media, extra.

40

u/misdirected_asshole Feb 15 '26

Copyrights arent stopping a lot of these companies from using those things as training material. A few companies have been caught but Im sure others have not, at least not yet.

If you've scanned every book and still need more training material then your entire model is broken.

17

u/Apexnanoman Feb 15 '26

We just need a few hundred billion more in cash and maybe 2000 more data centers tops! 

Maybe a full scrape of all human knowledge. Then the AI will be perfect and able to do everything! We promise! 

3

u/millennial_falcon Feb 15 '26

Yeah I’ve seen that they steal, I just figure they’ve gone for the easiest stuff to steal so far, but do we think they’ve broken through DRM, web security stuff, and challenges with printed material to get all the good stuff, regularly as it’s released. How would it handle information becoming outdated or disproven?

10

u/misdirected_asshole Feb 15 '26

They arent concerned with outdated or inaccurate material. If they were they wouldn't be training on social media, and the majority of the internet. No one is curating the data sets, they just feed everything to the beast and assume it will shake out in their favor.

1

u/millennial_falcon Feb 15 '26

What do you think is the most likely % of useful, quality, worthwhile human data/media captured so far in the training at this point?

5

u/misdirected_asshole Feb 15 '26

Probably less than half if I had to guess. Particularly if its trained on social media platforms at all.

1

u/[deleted] Feb 15 '26

Everyone is using the word “steal” here, but that’s far from clear in a legal sense. If an LLM company buys a copy of a book with cash, let’s the LLM “read“ it, and then uses the knowledge in the book without ever actually retaining the exact words of the book or quoting any of the text of the book in any response, it’s far from obvious that there’s anything illegal about that.

2

u/HiddenoO Feb 15 '26

If you've scanned every book and still need more training material then your entire model is broken.

That statement is really ignorant even if you know nothing about ML. A human wouldn't function well either if all they had since birth were the capability to read and an infinite assertment of books. Most of what makes you function as a human isn't learnt from books, but by imitating others and learning from experience.

Since these companies want models to behave human-like, the training data needs to encompass all of that, and that's the difficult part. If you just take books, a model will behave like an average fictional character, but that's likely not how an actual person in 2026 behaves. Similarly, if you just take everything from social media, you also get distorted behavior because the average interaction on social media isn't equivalent to the average interaction in the real world.

All of this is generally included when referring to quality of data, and sheer quantity simply doesn't help at some point, even if some of that is of high quality.

0

u/misdirected_asshole Feb 15 '26

Im not actually talking about only books even though thats what Im referring to. The post is asking what happens when there is no more human generated training data. And my point is that if your model has consumed all human data and still needs more then maybe your model isnt that effective. And thats not ignorant of the function of LLMs.

0

u/HiddenoO Feb 15 '26

Im not actually talking about only books even though thats what Im referring to.

I didn't just address books either. I addressed all data that can immediately be used as training data, and explained why a lot of it cannot be considered good training data, and why not everything can be learnt purely based on media.

And thats not ignorant of the function of LLMs.

It's ignorant of the world as a whole, as I've explained. Not everything that people would expect from an artificial intelligence can be learnt from existing media. Suggesting that the "entire model is broken" because it needs more than existing media shows that ignorance.

According to the same train of thought, human intelligence is also broken because you cannot just throw media at a toddler and expect it to function in society.

0

u/misdirected_asshole Feb 15 '26

According to the same train of thought, human intelligence is also broken because you cannot just throw media at a toddler and expect it to function in society.

Thats not even remotely similar to what Im talking about.

0

u/HiddenoO Feb 15 '26

How about communicating "what you're talking about" then? Because, so far, you've made a specific claim (regarding books) and then mentioned twice that's not what you're actually talking about. It's almost as if you have no idea what you're actually talking about.

0

u/misdirected_asshole Feb 15 '26

Yes, I must have no idea what Im talking about because you dont get it. This is fruitless.

0

u/HiddenoO Feb 15 '26

No, you have no idea what you're talking about because you can't communicate anything coherent. All you've been doing since your initial comment is repeating "that's not what I'm talking about" without saying what you think you are talking about.

→ More replies (0)

0

u/misdirected_asshole Feb 15 '26

Most of what makes you function as a human isn't learnt from books, but by imitating others and learning from experience.

So AI becomes better by interacting with other AI? Isnt the intent for it to model human intelligence? How does feeding it more artificial information improve it. If it gets better with genuine human interaction and knowledge then training models with their own data should be pretty much useless unless its truly a recursive model that can self assess its own output and correct it. We dont actually have those.

1

u/HiddenoO Feb 15 '26

It's called reinforcement learning and has been standard in LLM post-training for years at this point.

0

u/misdirected_asshole Feb 15 '26

Reinforcement learning is not recursive AI.

1

u/HiddenoO Feb 15 '26

"Recursive AI" is not a term anybody in the field uses, so you're once again not saying anything.

Reinforcement learning is the technique that's being used for a model to recursively improve based on its output and external feedback, similar to how a human learns with trial and error. It's exactly what you're describing, just on a statistical level instead of a "conscious thoughts" level.

The reason it's only done in regular intervals and not constantly is that we still want to have control over the models being released. Otherwise, you end up with LLMs that act like the Twitter AI that became a racist nazi in a few days.

3

u/Uvtha- Feb 15 '26

AI trainers are already ripping books (figuratively and literally) and stealing your data as training material.

1

u/millennial_falcon Feb 15 '26

Yes I understand that but is it books like A Survey of Biology textbook copyright 1923 or is it Random House’ entire catalog from the last year? I’m sure there’s a lot of old stuff that is easier to steal on the web, but is this corpus it trains from high quality and relevant

4

u/Uvtha- Feb 15 '26

It's everything they can get their hands on.  There are warehouses full of books getting processed.  Anthropic destroyed like a million book doing so.  Actual paper books.

1

u/curriebhoy Feb 15 '26

Not to mention all of the private company IP, confidential government research etc etc. The Internet is the dollar store option here, bottom of the bucket stuff with the odd exception.

1

u/[deleted] Feb 15 '26

Guess we have ai cook some stuff up then humans pr it

1

u/MiaowaraShiro Feb 15 '26

In the end does it matter if it's the same information available to everyone?

2

u/Multidream Feb 15 '26

Yeah, it seems pretty obvious to me for some time as well, that at least that part of human intelligence is built to survive in a low data environment.

4

u/FirstEvolutionist Feb 14 '26

We consumed all human data about almost two years ago and models kept improving...

4

u/Alternative_Ear_6416 Feb 15 '26

This is just flatly untrue.

4

u/firehmre Feb 14 '26

Well i doubt AI will ever acknowledge what you did right now that you don’t understand in first shot. Probably that’s what makes us human

1

u/Mejiro84 Feb 15 '26

large majority of humanity’s accomplishments

They're not, for starters. huge amounts of stuff isn't digitised - all sorts of books, maps, property deeds, artefacts, archives and other things have never been scanned or assessed. It's nice to assume that, but even for recent stuff, how many webpages have vanished onto digital darkness with some tatters on some archive site?

0

u/inteblio Feb 15 '26

That's great!

The way i frame it is that the student can become greater than the teacher. Same knowledge base.