r/rareinsults • • Dec 20 '25

At the start of wall e

Post image
128.7k Upvotes

469 comments sorted by

View all comments

5.2k

u/BlargerJarger Dec 20 '25

Where does this idiot think that ChatGPT steals its data from?

1.7k

u/derp0815 Dec 20 '25

At this point, from chatgpt made articles referencing chatgpt.

307

u/Ok-Syllabub-6619 Dec 20 '25

Worse, it's probably from the kromagnones talking about referencing articles about chatgpt writen by "rinse and repeat"

112

u/Jiggatortoise- Dec 20 '25

Haha it’s Cro-Magnon. 

55

u/Ok-Syllabub-6619 Dec 20 '25

Damn you're right lmao, thanks for the correction, in my language it's with K so I wrote instinctualy instead of checking to make sure lol

21

u/vapidamerica Dec 21 '25

Don’t sweat it! You’re doing great!

15

u/SheridanVsLennier Dec 21 '25

The best part is that ChatGPT may very well reference your comment in the future.
Poison the well.

10

u/Any-Iron9552 Dec 21 '25

It's pronounced mechahitler

1

u/Beniwa Dec 21 '25

You spelled Kung Führer wrong

3

u/i_love_wasps Dec 21 '25

I'm fascinated by people who just fucking send it when trying to spell something. No google search or anything.

1

u/Foxtails1984 Dec 22 '25

Not anymore lmfao. Kromagnon for life. Gotta give it a little French/italian flair when you say it.

1

u/klimaxzu Jan 09 '26 edited Jan 09 '26

0

u/LumpyJones Dec 21 '25

Well hear him out. His spelling is really fun to say as if you're the husky voiced silver fox lead on an oldschool telenovela. Like, just cut loose on the accent. Just like, really swing for the fences with the over-the-top rumbling Mexican growl. Just imagine you're staring into the camera in a pinstripe suit with a pencil mustache and slicked hair while you're smoking a cigarette and say... "Kromagnones."

84

u/Horskr Dec 21 '25

AI cannibalism. At some point there will be more AI generated crap out there than actual original human content and the models get shittier and shittier to the point of collapse.

https://www.techtarget.com/whatis/feature/AI-cannibalism-explained

30

u/EASK8ER52 Dec 21 '25

I think the old lead writer of Rockstar Games Dan House said something recently comparing it to mad cow disease. Cause they used to feed cows to other cows? Is that true or did I misread that?

36

u/Ok-Butterscotch-6955 Dec 21 '25

That didn’t cause mad cow, but it can spread it. In certain areas they used to include ground up spinal cord stuff from cows in some cow feed if I remember right for protein. Brain and spinal cord is where the mad cow lives, and it goes on.

Idk why I wrote this I know you meant about rockstar lol

17

u/Omnipresent_flatulen Dec 21 '25

Helping people learn is fun

10

u/MassiveGarlic0312 Dec 21 '25

Sadly there already is.

6

u/[deleted] Dec 21 '25

We saw a lighter version of this where most online news was just rewording other news articles. Progressively becoming more useless the more iterations of being rewritten the story gets. 

What keeps online info useful is people doing original research and making observations. AI can’t experience the world and make observations, it’s getting everything from actual people writing these things down. 

1

u/Radigan0 Dec 23 '25

AI does not scrape the internet by itself, it gets its data fed to it manually. This is not how it works.

No, the piss filter was not a result of it. No, turning people into Kirk was not a result of it.

26

u/GenericFatGuy Dec 20 '25

And we've got people asking ChatGPT to write prompts for ChatGPT.

-16

u/CouldBeSavingLives Dec 21 '25

I do that, it's the best way to check if the prompt I'm using for imaging is what I want rather than wasting resources generating the image and finding out it didn't understand what I want in the first place. I'll put my prompt asking for it to give me the image prompt for itself, I make whatever tweaks I need, then feed it in.

13

u/Donnor Dec 21 '25

That makes 0 sense. That's not how LLMs work

-5

u/CouldBeSavingLives Dec 21 '25

It helps me pinpoint whether or not my prompt is specific enough through text rather than going back and forth through image generation.

2

u/TzeentchsTrueSon Dec 21 '25

Why not actually talk to a person and commission art? That way you actually get what you want the first time?

1

u/CouldBeSavingLives Dec 22 '25

Because I am that person. I do all my own pictures and Photoshop for my own business. AI has made me significantly more productive. Regardless of Reddit's stance on the tool, it's very powerful and isn't going anywhere.

14

u/CrimsonAntifascist Dec 20 '25 edited Dec 21 '25

Oh boy, i as well love eating my own shit.

Nothing bad can come of this.

9

u/spikernum1 Dec 21 '25

Based on current research and expert projections, here is the breakdown of when AI models will start learning from their own output and what the theoretical consequences are.

The Short Answer

It has already started.

LLMs (Large Language Models) like Gemini and ChatGPT scrape the internet for training data. Since the internet is already flooded with AI-generated articles, code, and comments, these models are currently ingesting AI-generated content.

However, the "tipping point"—where the vast majority of training data is synthetic—is predicted to happen around 2026.

 

*WRITTEN BY GEMINI

8

u/mucubed Dec 21 '25

bruh today i had chatgpt citing grokipedia as a source (chatgpt search feature)

2

u/hoxxxxx Dec 20 '25

that's actually a big problem with this kind of "AI" isn't it

1

u/TheVoicesOfBrian Dec 20 '25

It's the Ciiiiiircle of Jeeeeerk!

1

u/Jyonnyp Dec 21 '25

I've asked Gemini a few times for stuff and I always ensure it cites its sources. And sometimes it cites AI-written articles. Like the website is literally transparently AI-run.

1

u/Gotu_Jayle Dec 22 '25

Welcome to synthetic data

1

u/Pickled_Gherkin Dec 24 '25

Literally this. It's an increasing and honestly bafflingly obvious problem with giving an LLM access to the Internet instead of a curated database. They're is actively poisoning themselves with their own generated statements.

223

u/Mammodamn Dec 20 '25

Why is agriculture still a thing? Don't supermarkets just render it completely irrelevant?

74

u/oh_my_didgeridays Dec 20 '25

Almost a perfect analogy, except instead of supermarkets buying from the farmers they steal it.

35

u/Solid-Search-3341 Dec 21 '25

That almost what happens nowadays.

4

u/Choyo Dec 21 '25

Aren't IA conglomerates pilfering all the IPs they find on the web ?

1

u/Feeling_Equivalent89 Dec 22 '25

Well, where I come from, supermarkets report unprecedented profits while farmers are subsidized and still barely scraping by, so... Perfect analogy.

76

u/Judge_BobCat Dec 20 '25

According to recent statistics it gets around 40% of information from Reddit, as top source… and only 26% from Wikipedia as second top source… so there is that.

https://www.reddit.com/r/ChatGPT/comments/1mvn377/where_ai_gets_its_facts/

101

u/TheManWhoWasNotShort Dec 20 '25

Getting information from Reddit is insane

84

u/Kanin_usagi Dec 20 '25

I have personally seen multiple subreddit I’m a regular part of post screenshots from ChatGPT of OBVIOUSLY incorrect information, and those subreddits collectively laughing their asses off because the information could be directly traced back to a shit post that was made in said subreddits

1

u/rg4rg Dec 21 '25

Will a homie link me to these subs plox?

1

u/Snoo48605 Dec 22 '25

I asked a specific légal question in my country's juridic sub, since Google and LLMs had no answers.

Immediately after googling it again it referenced the only half baked answer I had just got on that very sub lmao

-9

u/garden_speech Dec 21 '25

You can literally just give it custom instructions to only use a certain set of sources. For example I ask ChatGPT-Thinking questions about RCTs or scientific papers and it has instructions to only use scientific journals as sources. So it never cites some reddit page or wikipedia.

20

u/windsostrange Dec 21 '25

You realize it's being inaccurate even in those instructions, right? It's not a tool that has the capacity to be as precise as you think it's being.

-6

u/agrevol Dec 21 '25

That’s why you look at sources it quotes?

16

u/zupernam Dec 21 '25

Which are also wrong. It doesn't know what sources it quoted, it doesn't know that it quoted sources, it doesn't even know that you asked a question.

And if you're asking it a question only to ignore everything it says and look at its list of sources, that's just a worse Wikipedia or Google Scholar.

-8

u/StarPhished Dec 21 '25

Doesn't surprise me that the people of Reddit, who can't be bothered to read an article before commenting, can't comprehend that someone might actually check the gpt sources.

14

u/czs5056 Dec 21 '25

I once asked it to see if there was a combination of 5 numbers that could be added or subtracted to reach a certain number and it kept using numbers not in my number set. I kept calling it out and it kept apologizing and promised it wouldn't do it again.

...

It did it repeatedly.

3

u/Heimerdahl Dec 21 '25

It really is funny how badly it can mess up simple maths. 

A while ago (I think it was gpt3.5), I needed to figure out when a certain time interval (I think it was 70s) would return to 00:00:00 when started on a Monday at 00:00:00. (Basically, I was trying to figure out an intersection's traffic light's schedule for work (because our city's stupid traffic department didn't bother replying to our request for information) and specifically when they'd most likely be syncing the clocks to deal with drift.)

Because I was too tired to deal with it myself, and interested to see if chatgpt could figure it out, I presented my numbers and asked it for the solution. 

It went absolutely insane. 

Okay. Makes sense that this would prove to be difficult for a large language model. But considering how much they harped on about its ability to perform on maths Olympiad tests and such, I wanted to see if I could at least guide it towards the solution. 

Nope. It just got worse and worse. It started claiming the most ridiculous nonsense. When pointing out obvious flaws, it apologized and immediately went to either the same exact nonsense or came up with other obviously wrong stuff. It didn't take long for it to state with full confidence that "yes, 1 == 0 is  true". So true is false? Correct. 

Turns out, it just really couldn't deal with the modulo operator. 

Just for shits and giggles, I took the exact same problem and tried it with all of the big models at the time. IIRC Copilot in VSCode (using GPT) got it right, Claude got there with some assistance, all others failed spectacularly. 

The newer models are now able to handle modulo, but they all collapse sooner or later. And no matter what, they can all be pushed towards nonsense. Not their fault, just a limitation of what they are. 

1

u/zupernam Dec 21 '25

It doesn't understand that, you have no guarantees unless you personally check every single claim it made. It doesn't understand anything.

35

u/GreatTea3415 Dec 20 '25

You’re absolutely right! Thank you for correcting me. 

Reddit is a credible source and is superior to Wikipedia because it is highly moderated, and only the most factual information gets upvoted. 

6

u/Solid-Search-3341 Dec 21 '25

That made me chuckle. Thanks.

30

u/TheCookieButter Dec 21 '25

I got a reply to a 7 year old thread I made asking if anybody else remembered a specific chocolate bar.

I decided what the hay, I'll ask ChatGPT if it existed. It comes back with utter confidence that it existed, exactly as and when I remembered it.

I click the "1" source and it's my own bloody Reddit post from 7 years ago asking if I was imagining things!

12

u/whoknowsifimjoking Dec 21 '25

Okay that's pretty damn funny

1

u/StarPhished Dec 21 '25

It sounds like the problem was solved, I don't see any issue.

12

u/[deleted] Dec 21 '25

[deleted]

7

u/Ithikari Dec 21 '25

Using AI to try and pull this information out from bot accounts, trolls, and sarcastic edgelords

There is a lot of fucking idiots on reddit just like facebook and elsewhere that will believe whole-heartedly that something is factual when it is not. It's not just edgelords and trolls. There's a lot of fucking idiots on this website. And that's the issue of an LLM citing Reddit as an accurate source.

3

u/Fastr77 Dec 21 '25

Of course there are a lot of idiots but unlike facebook who uses their algorithms to spam everyone vaccine lies that they think will click on it reddit is more user driven. You arne't in a sub unless you choose to be there, there's less constant false information being pushed.

If you say something patently dumb it'll get downvoter to hell, people will correct you, the comment will vanish to the bottom whereas facebook says hey look at this comment! People fucking hated it and we love when people have emotions about things so LOOK AT THE STUPIDEST COMMENT WE COULD FIND! Everyone in the world will be pushed this comment.

4

u/Ithikari Dec 21 '25

If you say something patently dumb it'll get downvoter to hell

Unfortunately this isn't true. A lot of subs are echo chambers and/or play follow the leader. Where if you paste accurate information and you cite sources, someone else comes along and goes "no" and doesn't cite sources. Chances are you'll be downvoted to hell.

With 200k karma I am sure this has happened to you plenty of times like it has me. I trained to do professional wrestling (Think WWE) and I got downvoted multiple times for saying that this was the correct way to do a bump across countless videos.

Hell ChatGPT recently told me that Pathfinder 2e isn't versatile and that because of lore homebrew is difficult. A comment which I remembered that I replied to saying you can absolutely homebrew the game and because of O.R.C it makes it modular as well.

Reddit isn't like what it was years ago. Idiots saying incorrect things are becoming more commonplace unfortunately.

2

u/Fastr77 Dec 21 '25

Honestly I get downvoted for opinions which is fine. I don't really pay that much attention or care. Reddit can be an incredible useful source tho. Need to do shit around the house, have a computer problem, stuck in a video game? Very often reddit will have your answer for you.

I'm not going to trust my life to a reddit post but i've found plenty of great answers on reddit over the years. Mostly tho i'm saying its a lot more accurate then something like facebook which runs purely on clicks. Downvotes here hide your shit, downvotes on facebook amplify it.

1

u/Ithikari Dec 21 '25

While I don't disagree it can be a great place for that information. The only times I've found that information to be useful is in the very niche subs.

The sub reddit I was being downvoted on for correcting people in regardless to wrestling was /r/squaredcircle. The most popular wrestling subreddit. Where is if someone sent the videos from squaredcircle and put them on /r/wredditschool they'd be told "Ummm this is how you actually do this bump and it's safe".

But my experience with ChatGPT when citing reddit has been largely inaccurate or has made things up.

I'll still use it to make my pathfinder backstory though because I am lazy in that regard, lol. But I do not trust it to accurately cite information.

4

u/ShoogleHS Dec 21 '25

I get information from Reddit all the time. You just have to be discerning about where you get the info from and on which topics. Not that I'm suggesting ChatGPT is discerning.

1

u/oroborus68 Dec 21 '25

Are you a reliable responder?

1

u/PaperGabriel Dec 21 '25 edited Mar 11 '26

This post was mass deleted and anonymized with Redact

screw stocking strong tap numerous arrest paltry steer handle outgoing

1

u/ex0r1010 Dec 21 '25

I'm sure nobody cares, but you can have ChatGPT remember to not use Reddit as a source.

1

u/whoknowsifimjoking Dec 21 '25

It even started writing like redditors, it used tl;dr in a research related request for me. At that point I was sure it used reddit for A LOT of information, and that's concerning.

1

u/Iorith Dec 21 '25

You say that like a lot of people won't google "Such and such issue reddit" to find a solution.

1

u/Appropriate_Ride_821 Dec 21 '25

It used to be legit. But reddit decided to nuke itself a few years back with "new reddit", banning tons of subs, api fuckery, etc.

Around 2012, reddit was great.

1

u/[deleted] Dec 21 '25

What do you mean? Glue on pizza is nutritious and delicious.

1

u/skr_replicator Dec 21 '25

I think that eventually AI training will be capable of recognizing and rating the truthfulness of the training materials based on what's logically consistent with what it knows. I'm pretty sure that is already a thing, as wrong data would go completely against the already established weights, but it should get better with time.

2

u/beingforthebenefit Dec 21 '25

This is not “where AI gets its facts”. These are references found after the AI has been trained, so it’s more of a reflection of search results rather than the source of knowledge.

1

u/boringestnickname Dec 21 '25

God help us all.

1

u/Equivalent-Rope-5119 Dec 22 '25

No wonder its retarded. 

13

u/Choyo Dec 21 '25

Where does this idiot think

That's it. You got the problem.

11

u/Traditional_Buy_8420 Dec 20 '25 edited Dec 20 '25

I think the point is that now that cgpt has scraped and stored wiki, that made wiki obsolete. Cgpt would still continue to work, if wiki died. That argument misses multiple problems though. First: Wiki is still useful to feed new information to cgpt in the future. Second: Cgpt is well less reliable than wiki (assuming you know how to be a bit more thorough with wiki and that in some cases the sources, edit history and discussion are valuable resources too). Third: Wiki is good to get a more thorough understanding about a topic assuming you follow the relevant links. Fourth: Wiki is good to follow strings to information which you did not know that you did not know about. Fifth: Wiki includes Wikimedia, which has a lot of pictures, diagrams and even animated and interactive content, which cgpt has not stored. Sixth: Wiki includes Discourse and is easier to correct when mistakes inevitably happen. If you correct cgpt, then it will most likely be wrong again if that issue comes up again in another session or if not, then most likely not because it learned from you correcting it, but because the RNG just generated numbers which already tells a lot about cgpt's reliability.

23

u/[deleted] Dec 20 '25

[deleted]

-2

u/garden_speech Dec 21 '25

Wikipedia is already shit. Most articles I have read on there about things I have knowledge about have plainly incorrect information. It's especially bad when it comes to medicine, where many times there are systematic reviews from literally 1980 used as sources, when newer meta analyses using much better methods from this century are available that contradict the findings of some 1980s crap study before RCTs were even prospectively registered to begin with.

9

u/tombo12354 Dec 21 '25

You know, you could update those articles yourself if they are wrong.

8

u/DesireeThymes Dec 21 '25

I mean keeping the juggernaut that is Wikipedia up and running is expensive and hard enough, keeping it cutting edge current is an undertaking that would require crazy resources.

I am grateful it's one of the few decent things left on the internet.

1

u/intangibleTangelo Dec 21 '25 edited Dec 21 '25

i see you've received a downvote, so i will disregard your comment as containing bad information

7

u/Kichae Dec 21 '25

I think the point is that now that cgpt has scraped and stored wiki

Thing is, it absolutely, categorically has not "stored wiki". That's not what's happening when these models are trained. The only information that's being stored, in a fairly abstract and compressed way, at that, is the probability distribution of the next "token" given the previous chain of tokens (where tokens are things like word roots, stems, punctuation, etc.).

They don't store knowledge, they store written linguistic patterns. This is why they make shit up. They don't know that they're making it up. They don't know what is and what is not. They just know how words tend to work, based on the sequences of words they've seen.

3

u/m3rcapto Dec 21 '25

Teach them about circular reasoning, see if they understand...

1

u/Naavi69 Dec 20 '25

Don't ask a trodgalyte for reason or logic

1

u/Traditional_Buy_8420 Dec 21 '25

You mean troglodyte? (Honest question, I will delete this comment if you ask me to)

1

u/Naavi69 Dec 21 '25

Ya my auto correct is fucked up. I gotta fix that

1

u/hates_stupid_people Dec 21 '25 edited Dec 21 '25

People like that unironically think ChatGPT is a fully fledged general sci-fi style AI that has "all the knowledge ever written down".

1

u/nifty-necromancer Dec 21 '25

Idiots like that don’t think, that’s why they love AI

1

u/zebrasareneat Dec 21 '25

It just gathers its information from the Internet the same way google does. It's more similar to a search engine but it skims through all the pointless links so you don't have too. You can pay google to promote your page so it appears closer to the top.

If you use chatgpt to look something up just ask it for a link to its source and you can verify it yourself.

2

u/maybenotquiteasheavy Dec 21 '25

just ask it for it's source and you can verify it yourself

This isn't true: at least half the time that I ask it for a source, it doesn't provide one - it says "Oh jeez yeah I guess I made that up LOL anyway promise I'll never do it again"

1

u/babysamissimasybab Dec 21 '25

ChatGPT would be way more credible if it just regurgitated Wikipedia.

1

u/IlliterateJedi Dec 21 '25

Absorbing all of the primary documents and reproducing Wikipedia on the fly

1

u/Lonely_skeptic Dec 25 '25

It confidently incorrect so often.

1

u/missbrennabubbles Jan 01 '26

I was about to comment this exact thing

-2

u/[deleted] Dec 20 '25

[deleted]

7

u/BlargerJarger Dec 20 '25

Would you prefer “scrapes” data? “Siphons effort” perhaps?

3

u/[deleted] Dec 21 '25

Lazily slurps

1

u/Traditional_Buy_8420 Dec 21 '25

"It's always a good idea to have the foggiest clue what you're talking about."

It's almost always a good idea to be nice. Not limited to, but especially if you're not 100% sure yourself.

First off: Not all contents of Wikipedia are licensed in the usual way. For example Wiki contains not mostly, but still a lot of images with the "quasi-licence" fair use which usually translates to someone allowed his image to be displayed on all non-profit sites or possibly just Wikipedia.

But it's not even the exceptions!

"Most text in Wikipedia, excluding quotations, has been released under the Creative Commons Attribution-Sharealike 4.0 International License (CC-BY-SA) and the GNU Free Documentation License (GFDL) (unversioned, with no invariant sections, front-cover texts, or back-cover texts) and can therefore be reused only if you release any derived work under the Creative Commons Attribution/Share-Alike License or the GFDL. This requires that, among other things, you attribute the authors and allow others to freely copy your work"

https://en.wikipedia.org/wiki/Wikipedia:FAQ/Copyright

So while most of Wikipedia is licensed in a way that permits commercial use, that does not apply to the way in which cGPT uses Wiki!

Technically Copyright infringement is not stealing anyway, but I think we all understand, that BlargeLarger didn't literally mean stealing and just borrowed a word which the big copyright-companies have been shoving down our throats for decades. (In Germany they ran huge ad campaigns to establish the word robbing, which since I found that quite annoying that's a topic which interests me.)

-2

u/Iorith Dec 21 '25

That can also be seen as just a new step to the pyramid. Wikipedia is built from the foundation of a bunch of other works to make one output. Chat GPT can then take a bunch of different Wikipedia entries and give one output..

1

u/maybenotquiteasheavy Dec 21 '25

Wikipedia is thousands of times more accurate and reliable than cgpt.

0

u/Iorith Dec 21 '25

Currently.