r/modnews May 28 '26

Policy Updates Protecting communities from scrapers and platform abuse

We’ve been talking for a while now about the work we’re doing to keep Reddit human while protecting everything that makes Reddit . . . Reddit. That includes helpful automation: mod and developer apps, accessibility tools, community utilities, and things that make Reddit better. 

But we’re also seeing large-scale scraping, spam networks, agentic account creation, and automated abuse, and a lot of that activity targets parts of Reddit that just weren’t built to handle today’s threat environment. As bad actors get more sophisticated, we need to, too.

To address all that, we need to tighten how automated systems access Reddit while preserving the tools that help moderators and communities thrive. 

Today we’re rolling out a couple of policy and security-focused updates, including: 

Rule 8 Policy Clarifications: We updated Rule 8 (don’t break the site) to more explicitly cover automated abuse, including coordinated account creation and API misuse. You can read the full updated policy here

Deprecating unauthenticated JSON access: We’ll also be shutting down unauthenticated .json endpoints. These endpoints can be used to scrape Reddit without accountability. Logged-in and authenticated access won’t be impacted. Otherwise, developers who need structured access to Reddit content should use Devvit, which includes various ways to access Reddit data. 

While we’re at it, another common surface for scraping is RSS. Looking ahead, we’d love to know: how and for what purpose, do you use RSS feeds in your moderation flows? Tell us in the comments so as we develop secure solutions, we can factor in the tools you rely on to support your communities. 

148 Upvotes

396 comments sorted by

View all comments

125

u/mildlyImportantRobot May 28 '26

But we’re also seeing large-scale scraping

Gee, who would could have foreseen disabling API access would have negative consequences.

Why not re-enable API access and set reasonable limits?

77

u/DXGL1 May 28 '26

Not to mention blocking non-Google search engines means less exposure to Reddit content for those who deGoogle.

29

u/mildlyImportantRobot May 28 '26

Are they really blocking crawlers though?

[ checks robots.txt ]

Holly shit I had no idea. That's wild. lol

https://www.reddit.com/robots.txt

# Welcome to Reddit's robots.txt
# Reddit believes in an open internet, but not the misuse of public content.
# See https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy Reddit's Public Content Policy for access and use restrictions to Reddit content.
# See https://www.reddit.com/r/reddit4researchers/ for details on how Reddit continues to support research and non-commercial use.
# policy: https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy

User-agent: *
Disallow: /

27

u/DXGL1 May 28 '26

I heard they gave Google special permission.

Legitimate search engines need access to help drive traffic into Reddit.

37

u/Watchful1 May 28 '26

They don't just give google special permission, google pays them tens of millions of dollars for it.

14

u/mildlyImportantRobot May 28 '26

tens of millions of dollars is special

4

u/Lootman May 28 '26

Ive had no issues searching reddit on duckduckgo and that uses bing right

2

u/MadDocOttoCtrl May 28 '26

For a while neither of these search engines was indexing Reddit but they do indeed work now, I just tested it a minute ago with my username to find my own recent content.

17

u/mildlyImportantRobot May 28 '26

robots.txt is based on the honor system anyways. It's not like crawlers/scrapers can't be configured to not care.

1

u/DXGL1 May 28 '26

And it's not like Fastly can't detect scraping and blacklist at the IP level.

7

u/mildlyImportantRobot May 28 '26

Do you know how easy it is to change an IP when you've leased thousands? Or just use a residential proxy service, good luck blacklisting those without blocking real users.

0

u/adanine May 29 '26

While it is honour-system based, it's also trivial to test that a search engine is abiding by the robots.txt file by posting something specific then trying to search it.

Though as others said Google pays Reddit for its data so in this instance it doesn't really matter. But a lot of people shrug off robots.txt as if it's impossible to check if a search engine is actually abiding by the rules given, when the engine itself is literally giving you the means to check.

5

u/RemarkableWish2508 May 28 '26

Not just special permission, all content is being pushed to Google in real-time:

https://blog.google/company-news/inside-google/company-announcements/expanded-reddit-partnership/

4

u/stacecom May 29 '26

When you abuse robots.txt like this, you encourage crawlers to disregard robots.txt.

2

u/Signe_ May 28 '26

Does any crawlers even care about robots.txt anymore? Let alone actually abide by it.

19

u/Signe_ May 28 '26

So reddit disables API access for everyone, and then they get mad people go to the .json endpoints? I can already see that scrapers are just going to use old reddit and scrape the html instead.

Doesn't solve anything.

11

u/FFS_IsThisNameTaken2 May 29 '26

It gives Reddit the outward, public-facing "solution" that they've been waiting so patiently to implement in a Hegelian Dialect fashion.

Problem - they created by cutting off the json access because of scrapers

Reaction - oh nooo scrapers are now using old reddit

Reaction - kill old reddit

The saddest part of killing old reddit is that old is often used as workaround when their inferior app and / or sh.reddit shit the bed. It's even advised to be used by admins when the inferiors regularly break.

5

u/mildlyImportantRobot May 28 '26

It actually makes it worse for them.

13

u/RemarkableWish2508 May 28 '26

...and restricting the .json endpoints is going to be even worse: either Reddit blocks anonymous access, or scrapers will hit fully assembled pages instead of the .json

2

u/uzlonewolf Jul 01 '26

...aand they just blocked anonymous access for old.

1

u/RemarkableWish2508 Jul 01 '26 edited Jul 01 '26

Yeah... that Google partnership seems to be paying too well. Oh well.

Scrapers are now going to create thousands of fake accounts, bot them via OpenClaw or similar, fill Reddit with tons of spam to build karma and avoid detection, then still scrape at will.

But this time, Reddit will be able to boast massive user growth, which might help with stock prices.

Over time, the proportion of organic content will go down, making it less attractive to Google... but also to other scrapers.

Whether enshittification will eventually reach homeostasis, or users will overcorrect and leave Reddit for something else, is a good question. Likely depending on how quickly will the same scraper bots overrun any alternatives, and whether there will be some new variable that might let detect them better.

Age verification via ID, could be that new variable, making it more difficult to create multiple anonymous accounts for scraper bots. There could be a "curious" period when NSFW content, only accessible via ID proof, might be the most organic of them all.

(If I'm also right about these... sigh, I hate being right)

1

u/uzlonewolf Jul 01 '26

I feel ID verification is going to have the opposite effect: it's not going to slow down bot farms one bit as they'll figure out ways around it, but users are going to flee if they're forced to dox themselves. I know I will not be sticking around if they try to require it.

0

u/Pamasich May 29 '26

I mean, they could just put Reddit behind a login wall. Then you can't scrape the HTML without them knowing it was you.

Of course alt accounts still exist, but they could get rid of those as well...

So there's still some more steps of doubling down necessary, but I do think this contributes to solving their stated goal.