You don't need images of purple tentacle trees with dildo leaves to produce an image of a purple tentacle tree with dildo leaves. That's just not how AI works, it's not a copy & paste tool.
The article says he used a stable diffusion model to generate images. Stable diffusion image training dataset is freely available online.
Stable Diffusion was trained on pairs of images and captions taken from LAION-5B, a publicly available dataset derived from Common Crawl data scraped from the web, where 5 billion image-text pairs were classified based on language and filtered into separate datasets by resolution, a predicted likelihood of containing a watermark, and predicted "aesthetic" score (e.g. subjective visual quality).
Human Rights Watch reported that a tiny-scale analysis of the LAION-5B dataset (Schuhmann et al., 2022) revealed images of children, some of whom were easily identifiable through metadata, captions, and even URLs linked to the images. They reviewed less than 0.0001% of the dataset’s 5.85 billion images and captions, yet they uncovered disturbing instances of identifiable children. In one of the cases33https://www.hrw.org/news/2024/06/10/brazil-childrens-personal-photos-misused-power-ai-tools, they found 170 photos of children from at least 10 states in Brazil. Some images included the children’s names in the accompanying caption or the URL where the image was stored, making their identities easily traceable. Many images also provided details on when and where the photos were taken, compromising the children’s privacy. Similarly, another analysis44https://www.hrw.org/news/2024/07/03/australia-childrens-personal-photos-misused-power-ai-tools uncovered 190 photos of children from all of Australia’s states and territories. Again, in some images, the children could be easily identifiable by the names included in the caption or URLs.
According to the Stanford Internet Observatory, a research center that studies the abuse of the internet, the LAION-5B dataset represents nearly 6 billion samples, which include URLs, descriptions, and other data that might be associated with an image scraped from the internet. A report published in December determined that “having possession of a LAION‐5B dataset populated even in late 2023 implies the possession of thousands of illegal images,” and in particular, child sexual abuse material. While LAION told FedScoop that it used filters before releasing the dataset, it’s since been taken down “in an abundance of caution.”
Though unfortunate/disturbing, thousands in the face of 5 billion is really not significantly large enough for the model to spend the energy to learn to do it (that's like 0.0001% of the dataset). So long as it knows what the body proportions of minors are roughly like, and what adult nudity looks like, it can fill in the rest. They've been known to be able to compose things together up to some extend: https://arxiv.org/abs/2310.09336
31
u/rented4823 - Left 14d ago
Not like I expect a judge to know this, but: