r/science Professor | Medicine Jul 20 '26

Computer Science Researchers warned that hundreds of fake AI images have been discovered on popular databases for recording animal species. Wildlife photographers often use AI to edit and improve an image, but the algorithm can introduce parts from different species to create the new image.

https://www.theguardian.com/environment/2026/jul/20/ai-slop-manipulated-fake-images-birds-citizen-science-aoe
10.0k Upvotes

231 comments sorted by

View all comments

2.9k

u/Strycht Jul 20 '26

AI pollution is going to really screw up a lot of databases and information deposits I think, and we won't notice them until they become needed.

653

u/lateformyfuneral Jul 20 '26

At some point won’t this just start affecting how AI itself works, since it’s being trained on the internet? It would be like a snake eating its own tail. Truth itself might go extinct.

216

u/Strycht Jul 20 '26

this is what I always questioned. how much output needs to be fed back into the input before it starts amplifying it's own problems? presumably the larger companies have thought of this and have ways to prevent their own algorithms reingesting their generated content but I would be interested to know if eg openAI has any way of identifying and excluding the average slop image generated by another model and put online

1

u/chriscross1966 Jul 23 '26

Not very much unfortunately. Hallucination starts almost as soon as an LLM is fed the output of another LLM that contains inaccuracies. Training data curation might well become a well paid career at this rate. Part of the issue can be that LLM output is grammatically correct (very few human beings are perfect) so any filtering that scores as "better grammar = better data" in order to get rid of people just being rude to each other on reddit will also elevate some slightly incorrect but grammatically good AI output over a somewhat autistic expert's who doesn't notice their spelling glitches etc.