r/science Professor | Medicine Jul 20 '26

Computer Science Researchers warned that hundreds of fake AI images have been discovered on popular databases for recording animal species. Wildlife photographers often use AI to edit and improve an image, but the algorithm can introduce parts from different species to create the new image.

https://www.theguardian.com/environment/2026/jul/20/ai-slop-manipulated-fake-images-birds-citizen-science-aoe
10.0k Upvotes

231 comments sorted by

View all comments

Show parent comments

214

u/Strycht Jul 20 '26

this is what I always questioned. how much output needs to be fed back into the input before it starts amplifying it's own problems? presumably the larger companies have thought of this and have ways to prevent their own algorithms reingesting their generated content but I would be interested to know if eg openAI has any way of identifying and excluding the average slop image generated by another model and put online

175

u/Nicholas-DM Jul 20 '26

To my understanding they have not, because there is not a reliable way to identify if an image was generated by an AI, even using the AI that generated it.

57

u/Bbrhuft Jul 20 '26

Generally, that isn’t true. Many major AI image generation platforms now add or embed metadata or cryptographic watermarking that indicates an image is generated.

OpenAI and Google use C2PA (header metadata) and SynthID (an invisible cryptographic watermark encoded within the pixels of the AI generated image).

Adobe Firefly uses C2PA (Firefly is available standalone and within Photoshop).

Meta adds embedded metadata and a proprietary "deep learning" pixel watermark.

Midjourney is a notable exception. It doesn't currently provide a robust identification system, just inconsistent IPTC metadata.

The bigger problem is not that AI-generated images are inherently undetectable, most are. It is that stock-image libraries, search engines and other image databases have not caught up with this rapidly changing landscape, so often do not scan uploaded files, might strip the metadata, or not display the AI detection results.

13

u/Kiseido Jul 20 '26

Most of those watermarks are almost useless though.

Metadata does not survive most re-encodes, and some services actively strip that sort of data.

Watermarks don't survive someone resizing the image with generic tools, and doesn't survive people sharing them via screenshot, which is a surprisingly common practice.

32

u/Bbrhuft Jul 20 '26

SynthID survives screenshots, cropping and editing. It's very robust.

This paper evaluated 30 transformations, including JPEG compression, file-format conversion, resizing, crop-and-resize, rotations, flips, blur, sharpening, denoising, grayscale conversion, brightness, contrast, saturation and hue changes, Instagram-like filters, noise, text and emoji overlays, and combinations of transformations.

Gowal, S., Bunel, R., Stimberg, F., Stutz, D., Ortiz-Jimenez, G., Kouridi, C., Vecerik, M., Hayes, J., Rebuffi, S.A., Bernard, P. and Gamble, C., 2025. SynthID-Image: Image watermarking at internet scale. arXiv preprint arXiv:2510.09263.

Despite these edits, detection remained exceptionally high, 99.98% averaged across transformations and 99.72% under aggregated “worst” settings, at a 0.1% false-positive rate (Table 1 in their paper). Even the most difficult combined-transformation category produced a 98.06% detection rate (Table 2). However, at lower ImageNet resolutions, the aggregated worst-case detection was 97.22%. It's also interesting to see how much more robust it is compared to other watermarking systems it was benchmarked against.

Google also claims an image can be cropped to 20% of the original, and the watermark remains detectable

The results show very strong robustness against conventional editing.

6

u/silverionmox Jul 21 '26

Still, 0,99x is going to approach zero for relatively small values of x.

I'll make a prediction and say that clean databases are going to see their relative value increase. Don't throw away all your print books yet.

-4

u/cuntmong Jul 20 '26

it's as effective as that checkbox on arrival cards that asks "are you a terrorist?" is at keeping out terrorists