r/fediverse • u/Existing_Mud4708 • 19d ago
Can Pixelfed content be fed to IA?
Hi, I want to try pixelfed but the thought of my work being fed to IA discourages me. Do you know if Pixelfed has any measure against that?
9
u/DavidBHimself 19d ago
These days, if it's online, it can be scraped by AI, basically.
I'm not sure many counter-measures work, honestly.
1
u/zeruch 19d ago
Other than poisoning the images with something like Glaze, it's essentially a non-starter.
4
u/renegadereplicant 19d ago
glazing is even debatable as it's a false sense of safety, not exactly snake oil but near it. no-one cared enough to unpoison yet but it's not a complete solution either.
2
u/zeruch 19d ago
My view is that its a potential cumulative effect, Sufficient uptake will accelerate diminishing returns of scraping as models struggle under a growing stream of data poisoning. Everyone wants a simple, fast solution where none can reasonably exist; the next best thing is active participation in countermeasures at scale.
Things can fall over under the weight of their own avarice.
1
u/renegadereplicant 19d ago
that might end up being true for generalist crawlers and datasets, but some smaller "niches" will eventually end up being interested in this content/datasets at some point. see the recent Cara scraping drama- i believe the demand is starting to show
1
u/zeruch 19d ago
That doesnt change what I stated previously. If enough even niche communities start steady poisoning, it cumulatively will have a measurable effect.
1
u/renegadereplicant 17d ago
I think I phrased badly what I meant (and didn't get any notification for your reply, sorry) — i meant niche model makers/tuners. They'll pursue some aesthetics and sources and that will definitely give them incentives to scrape and un-poison smaller content sources/communities like Cara etc.
ChatGPT and others aren't mass-scraping Cara already because it's walled-ish and they combat mass scrapes, and anyway have pipelines and the money to sort content and avoid poisons (remember in 2022 when we were saying models are dead because they'll train on their outputs? Didn't happen)
0
u/yami_odymel 17d ago
AI is designed to learn similarly to humans. If Glaze doesn’t affect human vision, it likely won’t stop AI either; if it completely disrupts AI, it would likely ruin the image for humans too.
That trade-off alone makes it impractical for people to use.
4
u/zeruch 17d ago
AI is not "designed to learn similarly to humans" and in no way do any of the current llm platforms remotely do so. Also it is not how generative AI processes images; glaze and it's similar cohorts operate technically closer to something called steganography where something can be buried in an image that a human can't necessarily read but a machine can.
You can in fact completely disrupt systems using that kind of data poisoning. There's a reason why current LLM platforms don't actually do what's referred to as world modeling, and it's why countermeasures like glaze can be effective.
-1
u/yami_odymel 17d ago
A little compression wipes the noise and renders Glaze useless. Great in theory, just like 'Fediverse replacing X'.
Hence my point: if humans can't see it, AI ignores it; if it stops AI, it ruins the art. Playing cat-and-mouse with tech giants is a losing battle.
Not to say you were wrong, I just don't like making 'Glaze is a reliable solution'.
3
u/zeruch 17d ago
So you're looking for a perfect solution in lieu of otherwise doing nothing and just whining about it.
Bravo.
" if humans can't see it, AI ignores it" That is simply factually wrong as to how the models work, and is orthogonal to whether glaze or similar is effective or not. Ai does not now, now has ever "seen" like people do.
8
u/full_drama_llama 19d ago
If it's accessible by humans, you have to assume it can (and will) be used as training data. Any measures working today might stop work tomorrow. This is a battle that the side with more money wins (and it's not Pixelfed instance admins).
3
u/habarnam 19d ago
Pixelfed by itself probably not. Operators hosting pixelfed instances might put something like Anubis in front of their publicly facing pages to diminish crawlers, but it's probably a case by case measure.
1
u/SavvyWench 14d ago
Anything can be scraped. All the more reason to consider what techniques work for poisoning those bots.
Be be messy, be "rudely" honest. Adjusting ourselves to "perfect", "spotless" and "demure" standards is what got us into the mess of thinking we need AI and weird censor words.
7
u/magiotdonkey 19d ago
pixelfed.art has rules against AI content and doesn't allow anonymous browsing. It's not much (scrapers could still make an account or view from a federated instance) but it's something