Can AI crawlers actually read your site?

Check 36 AI crawlers against your robots.txt, then send a real request as each of the 15 whose user-agent its operator publishes. That second part is what catches a Cloudflare rule your robots.txt knows nothing about.

Frequently asked questions

What this check reads, what it sends, and what the result means

It reads your robots.txt and works out, for each of 36 AI crawlers and robots.txt control tokens, whether it is allowed or blocked, naming the exact line that decided it.

It then loads your homepage 15 more times, once as each crawler whose user-agent its operator publishes, which catches firewall and CDN blocks that robots.txt says nothing about. Those requests go out in small batches so your server sees a trickle rather than a spike.

See every crawler we check

Because robots.txt is a request, not a mechanism.

Cloudflare rules, WAF settings, rate limits and country blocks sit in front of it and answer first. A site whose robots.txt welcomes GPTBot and whose edge returns 403 to it is blocked in every way that matters, and only a real request reveals that.

Then your server refused us, and we report that instead of guessing.

A 403 or a bot-protection challenge on robots.txt is not the same as having no robots.txt, so the crawler rules stay marked unknown rather than green. We will not tell you every crawler is allowed on the strength of a file we never read.

It is also a finding in its own right: many AI crawlers fetch robots.txt from cloud addresses much like ours, so an edge rule that turns down our request may well be turning down theirs. The live request column still runs, so you can see whether your homepage answers a crawler user-agent even when the policy file does not.

No. Google-Extended is a setting, not a crawler.

Google crawls with Googlebot and consults the Google-Extended group only to decide whether that content may be used for Gemini training and grounding. Your Search presence and AI Overviews are governed by Googlebot, not by Google-Extended.

What Google-Extended controls

It depends which kind, and the three kinds do very different things.

  • Training crawlers: blocking keeps your content out of future models and changes nothing about today’s answers
  • Retrieval crawlers: blocking removes you from the indexes assistants cite right now
  • User-triggered fetchers: blocking fails a visitor who explicitly pasted your URL

Most sites that block all three did so with one wildcard rule and meant only the first.

What blocking each one costs

Yes, with a free Seerly account. No card, and no limit on checks.

Results are cached for a few hours so a site being checked repeatedly is only crawled once, and nothing is stored beyond that.

Look up a
specific crawler

Every crawler we check has a reference page: what it does with your content, the exact robots.txt rules to allow or block it, and what blocking it actually costs.