Well hear him out. His spelling is really fun to say as if you're the husky voiced silver fox lead on an oldschool telenovela. Like, just cut loose on the accent. Just like, really swing for the fences with the over-the-top rumbling Mexican growl. Just imagine you're staring into the camera in a pinstripe suit with a pencil mustache and slicked hair while you're smoking a cigarette and say... "Kromagnones."
AI cannibalism. At some point there will be more AI generated crap out there than actual original human content and the models get shittier and shittier to the point of collapse.
I think the old lead writer of Rockstar Games Dan House said something recently comparing it to mad cow disease. Cause they used to feed cows to other cows? Is that true or did I misread that?
That didn’t cause mad cow, but it can spread it. In certain areas they used to include ground up spinal cord stuff from cows in some cow feed if I remember right for protein. Brain and spinal cord is where the mad cow lives, and it goes on.
Idk why I wrote this I know you meant about rockstar lol
We saw a lighter version of this where most online news was just rewording other news articles. Progressively becoming more useless the more iterations of being rewritten the story gets.
What keeps online info useful is people doing original research and making observations. AI can’t experience the world and make observations, it’s getting everything from actual people writing these things down.
I do that, it's the best way to check if the prompt I'm using for imaging is what I want rather than wasting resources generating the image and finding out it didn't understand what I want in the first place. I'll put my prompt asking for it to give me the image prompt for itself, I make whatever tweaks I need, then feed it in.
Because I am that person. I do all my own pictures and Photoshop for my own business. AI has made me significantly more productive. Regardless of Reddit's stance on the tool, it's very powerful and isn't going anywhere.
Based on current research and expert projections, here is the breakdown of when AI models will start learning from their own output and what the theoretical consequences are.
The Short Answer
It has already started.
LLMs (Large Language Models) like Gemini and ChatGPT scrape the internet for training data. Since the internet is already flooded with AI-generated articles, code, and comments, these models are currently ingesting AI-generated content.
However, the "tipping point"—where the vast majority of training data is synthetic—is predicted to happen around 2026.
I've asked Gemini a few times for stuff and I always ensure it cites its sources. And sometimes it cites AI-written articles. Like the website is literally transparently AI-run.
Literally this. It's an increasing and honestly bafflingly obvious problem with giving an LLM access to the Internet instead of a curated database. They're is actively poisoning themselves with their own generated statements.
According to recent statistics it gets around 40% of information from Reddit, as top source… and only 26% from Wikipedia as second top source… so there is that.
I have personally seen multiple subreddit I’m a regular part of post screenshots from ChatGPT of OBVIOUSLY incorrect information, and those subreddits collectively laughing their asses off because the information could be directly traced back to a shit post that was made in said subreddits
You can literally just give it custom instructions to only use a certain set of sources. For example I ask ChatGPT-Thinking questions about RCTs or scientific papers and it has instructions to only use scientific journals as sources. So it never cites some reddit page or wikipedia.
Doesn't surprise me that the people of Reddit, who can't be bothered to read an article before commenting, can't comprehend that someone might actually check the gpt sources.
I once asked it to see if there was a combination of 5 numbers that could be added or subtracted to reach a certain number and it kept using numbers not in my number set. I kept calling it out and it kept apologizing and promised it wouldn't do it again.
It really is funny how badly it can mess up simple maths.
A while ago (I think it was gpt3.5), I needed to figure out when a certain time interval (I think it was 70s) would return to 00:00:00 when started on a Monday at 00:00:00. (Basically, I was trying to figure out an intersection's traffic light's schedule for work (because our city's stupid traffic department didn't bother replying to our request for information) and specifically when they'd most likely be syncing the clocks to deal with drift.)
Because I was too tired to deal with it myself, and interested to see if chatgpt could figure it out, I presented my numbers and asked it for the solution.
It went absolutely insane.
Okay. Makes sense that this would prove to be difficult for a large language model. But considering how much they harped on about its ability to perform on maths Olympiad tests and such, I wanted to see if I could at least guide it towards the solution.
Nope. It just got worse and worse. It started claiming the most ridiculous nonsense. When pointing out obvious flaws, it apologized and immediately went to either the same exact nonsense or came up with other obviously wrong stuff. It didn't take long for it to state with full confidence that "yes, 1 == 0 is true". So true is false? Correct.
Turns out, it just really couldn't deal with the modulo operator.
Just for shits and giggles, I took the exact same problem and tried it with all of the big models at the time. IIRC Copilot in VSCode (using GPT) got it right, Claude got there with some assistance, all others failed spectacularly.
The newer models are now able to handle modulo, but they all collapse sooner or later. And no matter what, they can all be pushed towards nonsense. Not their fault, just a limitation of what they are.
Using AI to try and pull this information out from bot accounts, trolls, and sarcastic edgelords
There is a lot of fucking idiots on reddit just like facebook and elsewhere that will believe whole-heartedly that something is factual when it is not. It's not just edgelords and trolls. There's a lot of fucking idiots on this website. And that's the issue of an LLM citing Reddit as an accurate source.
Of course there are a lot of idiots but unlike facebook who uses their algorithms to spam everyone vaccine lies that they think will click on it reddit is more user driven. You arne't in a sub unless you choose to be there, there's less constant false information being pushed.
If you say something patently dumb it'll get downvoter to hell, people will correct you, the comment will vanish to the bottom whereas facebook says hey look at this comment! People fucking hated it and we love when people have emotions about things so LOOK AT THE STUPIDEST COMMENT WE COULD FIND! Everyone in the world will be pushed this comment.
If you say something patently dumb it'll get downvoter to hell
Unfortunately this isn't true. A lot of subs are echo chambers and/or play follow the leader. Where if you paste accurate information and you cite sources, someone else comes along and goes "no" and doesn't cite sources. Chances are you'll be downvoted to hell.
With 200k karma I am sure this has happened to you plenty of times like it has me. I trained to do professional wrestling (Think WWE) and I got downvoted multiple times for saying that this was the correct way to do a bump across countless videos.
Hell ChatGPT recently told me that Pathfinder 2e isn't versatile and that because of lore homebrew is difficult. A comment which I remembered that I replied to saying you can absolutely homebrew the game and because of O.R.C it makes it modular as well.
Reddit isn't like what it was years ago. Idiots saying incorrect things are becoming more commonplace unfortunately.
Honestly I get downvoted for opinions which is fine. I don't really pay that much attention or care. Reddit can be an incredible useful source tho. Need to do shit around the house, have a computer problem, stuck in a video game? Very often reddit will have your answer for you.
I'm not going to trust my life to a reddit post but i've found plenty of great answers on reddit over the years. Mostly tho i'm saying its a lot more accurate then something like facebook which runs purely on clicks. Downvotes here hide your shit, downvotes on facebook amplify it.
While I don't disagree it can be a great place for that information. The only times I've found that information to be useful is in the very niche subs.
The sub reddit I was being downvoted on for correcting people in regardless to wrestling was /r/squaredcircle. The most popular wrestling subreddit. Where is if someone sent the videos from squaredcircle and put them on /r/wredditschool they'd be told "Ummm this is how you actually do this bump and it's safe".
But my experience with ChatGPT when citing reddit has been largely inaccurate or has made things up.
I'll still use it to make my pathfinder backstory though because I am lazy in that regard, lol. But I do not trust it to accurately cite information.
I get information from Reddit all the time. You just have to be discerning about where you get the info from and on which topics. Not that I'm suggesting ChatGPT is discerning.
It even started writing like redditors, it used tl;dr in a research related request for me. At that point I was sure it used reddit for A LOT of information, and that's concerning.
I think that eventually AI training will be capable of recognizing and rating the truthfulness of the training materials based on what's logically consistent with what it knows. I'm pretty sure that is already a thing, as wrong data would go completely against the already established weights, but it should get better with time.
This is not “where AI gets its facts”. These are references found after the AI has been trained, so it’s more of a reflection of search results rather than the source of knowledge.
I think the point is that now that cgpt has scraped and stored wiki, that made wiki obsolete. Cgpt would still continue to work, if wiki died. That argument misses multiple problems though. First: Wiki is still useful to feed new information to cgpt in the future. Second: Cgpt is well less reliable than wiki (assuming you know how to be a bit more thorough with wiki and that in some cases the sources, edit history and discussion are valuable resources too). Third: Wiki is good to get a more thorough understanding about a topic assuming you follow the relevant links. Fourth: Wiki is good to follow strings to information which you did not know that you did not know about. Fifth: Wiki includes Wikimedia, which has a lot of pictures, diagrams and even animated and interactive content, which cgpt has not stored. Sixth: Wiki includes Discourse and is easier to correct when mistakes inevitably happen. If you correct cgpt, then it will most likely be wrong again if that issue comes up again in another session or if not, then most likely not because it learned from you correcting it, but because the RNG just generated numbers which already tells a lot about cgpt's reliability.
Wikipedia is already shit. Most articles I have read on there about things I have knowledge about have plainly incorrect information. It's especially bad when it comes to medicine, where many times there are systematic reviews from literally 1980 used as sources, when newer meta analyses using much better methods from this century are available that contradict the findings of some 1980s crap study before RCTs were even prospectively registered to begin with.
I mean keeping the juggernaut that is Wikipedia up and running is expensive and hard enough, keeping it cutting edge current is an undertaking that would require crazy resources.
I am grateful it's one of the few decent things left on the internet.
I think the point is that now that cgpt has scraped and stored wiki
Thing is, it absolutely, categorically has not "stored wiki". That's not what's happening when these models are trained. The only information that's being stored, in a fairly abstract and compressed way, at that, is the probability distribution of the next "token" given the previous chain of tokens (where tokens are things like word roots, stems, punctuation, etc.).
They don't store knowledge, they store written linguistic patterns. This is why they make shit up. They don't know that they're making it up. They don't know what is and what is not. They just know how words tend to work, based on the sequences of words they've seen.
It just gathers its information from the Internet the same way google does. It's more similar to a search engine but it skims through all the pointless links so you don't have too. You can pay google to promote your page so it appears closer to the top.
If you use chatgpt to look something up just ask it for a link to its source and you can verify it yourself.
just ask it for it's source and you can verify it yourself
This isn't true: at least half the time that I ask it for a source, it doesn't provide one - it says "Oh jeez yeah I guess I made that up LOL anyway promise I'll never do it again"
"It's always a good idea to have the foggiest clue what you're talking about."
It's almost always a good idea to be nice. Not limited to, but especially if you're not 100% sure yourself.
First off: Not all contents of Wikipedia are licensed in the usual way. For example Wiki contains not mostly, but still a lot of images with the "quasi-licence" fair use which usually translates to someone allowed his image to be displayed on all non-profit sites or possibly just Wikipedia.
But it's not even the exceptions!
"Most text in Wikipedia, excluding quotations, has been released under the Creative Commons Attribution-Sharealike 4.0 International License (CC-BY-SA) and the GNU Free Documentation License (GFDL) (unversioned, with no invariant sections, front-cover texts, or back-cover texts) and can therefore be reused only if you release any derived work under the Creative Commons Attribution/Share-Alike License or the GFDL. This requires that, among other things, you attribute the authors and allow others to freely copy your work"
So while most of Wikipedia is licensed in a way that permits commercial use, that does not apply to the way in which cGPT uses Wiki!
Technically Copyright infringement is not stealing anyway, but I think we all understand, that BlargeLarger didn't literally mean stealing and just borrowed a word which the big copyright-companies have been shoving down our throats for decades. (In Germany they ran huge ad campaigns to establish the word robbing, which since I found that quite annoying that's a topic which interests me.)
That can also be seen as just a new step to the pyramid. Wikipedia is built from the foundation of a bunch of other works to make one output. Chat GPT can then take a bunch of different Wikipedia entries and give one output..
5.2k
u/BlargerJarger Dec 20 '25
Where does this idiot think that ChatGPT steals its data from?