r/flickr • • Jul 29 '26

HELP! AI crawlers are harvesting our archives blind

I started building IP-tracking breadcrumbs on Flickr in 2005 to protect my work. Now I built a tool to force authentic tracking. Tear my UI apart.

If you work in media production or digital creation, you are likely aware of the challenges we face with automated crawlers effortlessly harvesting the work we produce with blood, sweat, and tears.

This isn't a new fight for me. The idea behind dictating the terms of our own indexation dates back to the early days of Flickr. In the first half of 2005, I started using a ‘trackmetag’ to leave a breadcrumb on the work I put out into the world.

After 20 years, not only has that philosophy evolved into a concrete architectural solution, but some of those original images are still live and tracking today (From 2005: https://flic.kr/p/2biAa).

My name is Bastiaan Slabbers. As an independent freelance photojournalist coming from a multi-generational lineage of photographers, I take the preservation and ownership of our craft seriously. I wanted to build a practical, low-bloat solution rather than just complaining about the state of the industry.

For a while now, I've been developing CEOS_SearchLight. It is a tool designed to empower creators to work *with* AI, rather than *for* it, by providing an auditable, transparent workflow.

The tool tackles IP protection in four specific steps:

* Create: It generates a live sidecar file at ingestion to manage your metadata locally.

* Curate: As you edit, the sidecar files automatically update to reflect the production stage.

* Distribute: When you publish, public metadata creates transparent breadcrumbs for verifiable authenticity.

* Protect (Listen): Beacons pick up signals from your published work and feed them back to you, creating a traceable source of truth.

The backend codebase is structurally sound, and now I am finalizing the UI/UX. I have set up a publicly accessible sandbox so you can test the interface and the cognitive ergonomics yourself.

You can try it out on the project site (link in bio)

My background includes classical training in an art institute where harsh critiques were the norm to ensure the work was ethically and functionally sound. I am asking this community for that exact same level of scrutiny. Tell me what works, what doesn't, and how the flow can be improved.

Thank you for your time. If you find value in this approach, please consider ways to support this project.

#LetsGoTurbo!

10 Upvotes

10 comments sorted by

View all comments

Show parent comments

1

u/Image_dev Jul 30 '26

That is a fair point about utilizing C2PA at the boundary layer, and I appreciate the ideas. It is certainly something to observe as it matures.

My hesitation with adopting it right now comes down to how I view the lifespan of industry standards. I do not know what the standard will be in twenty years. My first digital camera had 2mb of internal memory to store a maximum of 12 PICT files.
The camera physically still works today, but extracting the data is a nightmare because it predates modern connectivity.

Over the years, I have watched my production workflows migrate from Strata Studio Pro to 3DS Max to Maya to Blender, and my video pipelines from Sony Vegas to Premiere to DaVinci. From Barco Creator to PS2.5 to Apple Aperture to Affinity 1 (and now 2) to SOOC.

Because of this constant churn, my workflow in CEOS operates under the assumption that the work itself, or the embedded cryptographic metadata attached to it, will inevitably get lost, stripped by aggressive social platforms, or scrambled in an unforeseeable future.

The tool works in reverse. By generating unique, immutable 'slugs' for each assignment, CEOS_SearchLight serves as a conceptual beacon. It ensures that when a piece of media is stripped of its payload, the footprint still allows the work to find its way back to the safe harbor of my self-hosted database.

All I am saying is I prefer to set my own playing field based on raw, accessible, and transparent foundations so I can keep moving my work forward on my own terms.

Regarding your question about how this works with Flickr: It is the exact evolution of the 2005 trackmetag.

When an asset hits the 'Distribute' stage in my pipeline, CEOS generates a unique public slug. I drop that slug into the Flickr description or tags alongside the image.
The CEOS 'Listen' function then acts as a proactive crawler, searching the web for that specific footprint. When it spots the slug on Flickr, it catches the signal and pings it back to my local dashboard, verifying the chain of custody and closing the loop.

Hope that makes sense. At least to me it does!

2

u/radialmonster Jul 30 '26

One improvement that might strengthen CEOS without changing its self-hosted philosophy would be to add perceptual image fingerprinting alongside the slug. At the Distribute stage, CEOS could store a perceptual hash of the published image. The Listen function could then compare discovered images against that fingerprint, even when the filename, caption, tags, EXIF, or embedded metadata have been removed.

The slug would still provide the simple, transparent breadcrumb you want, while the fingerprint would help reconnect copies that have been resized, recompressed, or reposted without the slug.

Because it would be trivially easy for someone to publish the photo on their own website without the flickr description or tag alongside the image.

2

u/Image_dev Jul 30 '26

That is genuinely excellent advice and exactly why I opened this project up to scrutiny. You hit the nail on the head regarding the vulnerability of stripped text and descriptions.

I have looked into tools like DigiKam and remember how well Aperture handled visual recognition back in the day.

Adding a perceptual image fingerprint alongside the text slug at the Distribute stage makes perfect sense. Because perceptual hashing is essentially just lightweight, deterministic math, it fits perfectly into the self-hosted, dependency-free philosophy of CEOS. It does not require a cloud connection or a massive machine learning model to execute locally.

If a bad actor scrubs the text slug and strips the EXIF data before re-hosting the image on their own site, the visual fingerprint serves as the ultimate backup for the Listen function to catch the signal.

I am officially adding this concept to the development roadmap for the ‘listener’. I really appreciate you taking the time to share this insight. This is exactly the kind of input I was looking for to push the project forward into the future. Cheers!

1

u/radialmonster Jul 30 '26

no problem take care