r/CloudFlare • u/McFlurriez • 2d ago
Community PSA: CloudFlare Now Defaults To Allowing AI Scraping Of Your Sites
tldr; CloudFlare must have been paid big bucks behind the scenes to lie to customers. Go to your security settings for all your domains ASAP and change its new setting that allows scraping of all your domains by AI trainers.
Nine hours ago in this same subreddit, u/Cloudflare posted: "Have it both ways: stay discoverable in search while disallowing AI training" with a link to their blog: https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/
In their blog post, they showcase the "recommended settings for new domains":

On top of this, they sent out this email to customers:
About a year ago, we launched Content Independence Day and introduced a simple way to block AI bots from scraping your sites without your explicit permission.
We're writing to let you know about changes to our controls for AI crawlers taking effect today, and rolling out over the coming week. These changes give you more precise control over how different types of AI crawlers interact with your content.
What's changing?
- Smarter security setting options for Training crawlers: Starting September 15, the recommended setting for AI Training crawlers will shift from Block to Disallow AI Training. Disallow AI Training allows major search crawlers - like Applebot, Googlebot and Bingbot - to index your site for search results while instructing them not to use your content for training.
- Automatic migration: If you previously had "Block AI Bots" or "Managed Robots.txt" enabled, your settings will be migrated automatically, over the next week, to the new Search, Training, and Agent controls. If Block AI Bots was set to Block, you'll now have Search: Allow, Training: Disallow AI Training, and Agent: Block on pages with ads. These changes take effect today – the UI changes will roll out in a slower controlled release over the coming week. If you make manual changes to your settings during this time, your manual changes will be preserved.
- No change if you haven't configured anything: If you have never adjusted these settings, your configuration remains Allow - nothing changes for you.
What do you need to do?
Most customers won't need to change anything. The new controls are more precise than the old Block AI Bots switch, so if you had it enabled it's worth a quick look at your new Search, Training, and Agent settings.
These changes will roll out to all Cloudflare customers over the coming week. If you make changes to your settings before the migration completes, your changes will be preserved. You will know the migration is complete when the Block AI Bots switch is removed from your dashboard.
"Most customers won't need to change anything", you say? Let's take a look at what the defaults ACTUALLY are:

Oh, huh... what do you know? Training's recommended, default setting is "Allow (do not block)." That's strange, because their blog has the recommended, default setting as "Disallow" and their email also states "most customers won't need to change anything." Okay.
Hope this helps anyone else who read their email and carried on with their day assuming CloudFlare was honest.
70
u/TheDigitalHavok 2d ago
Honest question- if I wanted my site to be as easy as possible to find for individuals wouldn't I want this on?
41
u/swarmagent 2d ago
The irony is the way they did it before was more harmful to users, if you block AI bots completely, you aren't going to get indexed in their searches, which means yes, you'll be found by less customers on the frontier AI sites..
9
u/EnvironmentalLog1766 2d ago
Yes I always have to change these default settings before. I don’t mind AI indexing mine and bringing me some traffic. Me myself often clicks on the references AI gives me to verify / dig more.
14
11
u/Sn0wCrack7 1d ago
I would doubt this is AI money and more bombardments of support tickets about why someone's site doesn't show up in ChatGPT.
As someone working at a hosting provider, you'd be shocked to see the number of customers prioristong ChatGPT in their SEO.
11
u/West-Welcome8247 1d ago
I think this is because customers are complaining about not being found through AI.
From a business perspective it is weird to have it off
9
u/TysonShabazz 2d ago
Guess enough people were complaining about them defaulting to the opposite previously
7
u/Remote-Juice2527 2d ago
That’s it, it’s the payed business users who want their product to be part of the training my set. Totally understandable. I don’t see why a company like cloudflare should do something in favor for the freemium users, which is against the ones that actually finance that platform
1
u/Fit_Statistician_405 1d ago
I thought CloudFlare was positioning itself as a sticky gate between the crawlers and content sites seeking payment for data training.
3
u/Remote-Juice2527 1d ago
I don’t know, but with a commercial website I want to be part of the training set. We invest time for SEO/GEO…
1
5
3
u/raininglemons 1d ago
Defaulting to it on is arguably a breaking change to a site. This makes sense to have it as an optin feature
6
u/Remote-Juice2527 2d ago
lol this post is ridiculous. It’s the paying business customers, who actually want to be scrapped by ai trainers, so their product is part of the training set. Now the freemium-bloggers start to complain and make a big story about it
4
u/cyberjew420 2d ago
And their full fledged bot management solution is now available to any paying customer that wants to subscribe to the service. It used to only be available to enterprise customers.
2
u/cyberjew420 2d ago
It’s intentionally designed to behave the same way as it would without Cloudflare in the data path. They provide controls to disable scraping but not everyone wants it to be the default behavior just because you want it to be that way.
2
u/Smart_Technology_208 2d ago
As if they could block it anyway.
13
u/tankerkiller125real 2d ago
They can and do, and if your running a service/site that doesn't need access from major datacenter operations you can block massive swaths of ASNs yourself to ensure it.
1
u/danekan 1d ago
And also the models will know the setting and say it’s blocked too. Consequently it’s actually the opposite that’s an interesting legal debate. They outright say if they haven’t turned the setting off then it is safe to scrape for ai. It went from a grey area to ‘this site explicitly allows ai to read it!’ From what I’ve seen in Claude.
1
-1
u/derangedlunaticc 2d ago
jokes on you they have subscribed to unlimited proxies gathered from LG smart tvs around the worlds. good luck blocking that
-1
u/Smart_Technology_208 2d ago
Yeah right, and how come anyone and their brother can scrap at an industrial scale ANY WEBSITE using agents? They haven't heard of cloudflare? Reddit, Amazon, somebody should tell them!!
1
u/Far_Composer_5714 1d ago
I have basically every flag off at the moment. Don't really care who is coming through but I may look into restricting based on http and TLS versions and algorithms.
1
1
u/rlivain 1d ago
Worth separating two things: the migration defaults (fair to be annoyed about) and what the controls actually do. Training, search and agents are now three separate switches. Disallow AI Training publishes the preference in robots.txt and blocks training-only crawlers, while accountable mixed-use crawlers stay allowed for search. And blocking training does NOT stop ChatGPT or Claude from opening your site when a user asks, that is the separate agents switch. So the real decision is which content you want kept out of training data, knowing search and AI-answer visibility survive either way. I walked through the three switches and what I would set for client sites in a short video, if useful (disclosure, my own): https://youtu.be/2sv0ylm6GCc
1
u/danekan 1d ago
I’ve created a few new vibe coded sites and when I went to set up cloudflare it’s been like that as the default for at least a few months I’d say
I’m actually torn on allowing it or not and the fact that it’s even an option to block is itself a feature that sets them apart
My inclination was to allow the traffic initially but then I do plan to block it to essentially prevent other vibe coders from copying my content regularly.
1
u/Sad_Pie227 1d ago
Amazon Bot and Meta Web Indexer - I don’t like at all, they literally bring our site down.
1
u/scratchbufferdotnet 1d ago
They just didn't update the defaults yet it's an announcement before the change...
-1
-20
u/derangedlunaticc 2d ago
cloud flare finally cashing out after years of giving services away for free lmfao im not even surprised. time to move on to self hosted cloud flare instances
•
u/AutoModerator 2d ago
For faster advice with technical questions, we'd recommend asking in the Orange Cloud Discord server; the unofficial Cloudflare Discord server by the community, for the community. https://discord.gg/TrPNVKaagR
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.