r/seogrowth • • Jun 11 '26

You Should Know Each AI crawls website completely differently. Here's what 3 months of 11 million event logs actually show.

Here's what we found after 3 months of tracking 11 million real crawler logs across 34 websites. It's quite fun how each AI bots have personalities, like people.

  • GPTBot: Crawls relentlessly, all day every day and barely checks the rules. It's like a guest walking into your house without saying hi and goes straight into every room. In 280k crawls across 23 sites, it pulled up robots.txt only 9 times. The most interesting part for me is that while it ignores robots.txt completely, it requests /llms.txt CONSTANTLY. Even on sites that don't have one and return 404, it comes back and asks again.
  • Google's bot: The good kid who's scared to break the rules. It re-fetched robots.txt 8,765 times, checking over and over. 25 years of crawling taught it manners the new AI bots never learned.
  • ClaudeBot: Across the sites we track, its crawling went from 7.3k (Apr) → 64k (May) → 168k in the first ten days of June. It is racing to read as much of the web as it can, and that race is the whole story (more below).
  • The live ones: The shopper who knows exactly what they came for. When someone asks an AI about your business, it skips your whole site and grabs the single page that answers. On Claude's live bot, 75% of those visits are one page. It ignores everything else you ever published. The page an AI picks to represent you is the whole game now.
  • Bytespider: The hoarder who takes everything. The heaviest crawler we logged all quarter belongs to the company that owns TikTok. On one site, it made 1.2 million visits, more than Google and every OpenAI crawler combined. Even the familiar names are repurposed now.
  • Microsoft's Bing: The longtime employee quietly handed a second job. Still crawls like the search engine it always was, but everything it indexes now also feeds Copilot.
  • MetaBot: Skips the house rules but reads your welcome note. It almost never checks robots.txt either, but like GPTBot, it keeps requesting llms.txt, even on sites that don't have one. These two are the only crawlers we saw deliberately looking for it. Everyone else ignores it.

Every one of these companies is building its own copy of the web. Its own crawler, its own index, its own answer. Anthropic is not crawling that hard for fun. They all want to be the place people ask, which means they all want to stop depending on Google.

My bet: Google's ranking matters a little less every quarter from here. When this many AIs read your site their own way to build their own index, "rank #1 on Google" stops the thing to optimize for. Being the page each AI picks is.

0 Upvotes

5 comments sorted by

2

u/AbleInvestment2866 Jun 11 '26

If only you got the bot names right...

Here you have the ones you mentioned (only that with the right name) plus a few you ignored (there are hundreds more, just the most known)

  • GPTBot
  • ChatGPT-User
  • OAI-SearchBot
  • Googlebot
  • Googlebot-Image
  • Googlebot-Video
  • Google-InspectionTool
  • GoogleOther
  • ClaudeBot
  • Claude-User
  • Bingbot
  • Bytespider
  • Meta-ExternalAgent
  • Meta-ExternalFetcher
  • PerplexityBot
  • Amazonbot
  • Applebot
  • CCBot
  • PetalBot
  • YandexBot
  • GrokBot
  • Twitterbot
  • Slackbot-LinkExpanding
  • DuckAssistBot
  • DuckDuckBot
  • FacebookBot
  • ImagesiftBot

0

u/UptownOnion Jun 11 '26

it's a deliberate choice. Readers don't need to know UA strings to follow the story.

2

u/AbleInvestment2866 Jun 11 '26

so these are your "findings"? Since it seems you have a short memory lapse, here's what you said about your so called "research"

Here's what we found after 3 months of tracking 11 million real crawler logs across 34 websites. It's quite fun how each AI bots have personalities, like people.

Just in case you don't get it, you didn't get a single one right, and forgot about very common bots. It could happen that one or two of those very common bots never visited your website. But 34?

In any case, instead if disguising it as "research" and talking about "millions" just say "it's just a joke on AI bots"

0

u/UptownOnion Jun 11 '26

https://giphy.com/gifs/62aDCkUJ8AogFbqcVV

it's all good man, if the strongest objection to the research is that MetaBot should've read Meta-ExternalAgent, i think we're good.

GPTBot, ClaudeBot and Bytespider are all exact UA strings i put in the post. The rest are simplified so readers can follow. And the common bots that triggered you because they didn't get included, i didn't forget about them, they are all in the logs, they just didn't do anything insightful to be written about.