Can AI crawlers actually read your site?
Check 36 AI crawlers against your robots.txt — and, for the six that matter most, send a real request carrying that crawler’s user-agent. That second part is what catches a Cloudflare rule your robots.txt knows nothing about.
Common questions
- What does the AI Crawler Access Checker do?
- It reads your robots.txt and works out, for each of 36 AI crawlers and robots.txt control tokens, whether it is allowed or blocked — naming the exact line that decided it. It then loads your homepage sending each major crawler’s user-agent, which catches firewall and CDN blocks that robots.txt says nothing about.
- Why does robots.txt say a crawler is allowed when the check says it is blocked?
- robots.txt is a request, not a mechanism. Cloudflare rules, WAF settings, rate limits and country blocks sit in front of it and answer first. A site whose robots.txt welcomes GPTBot and whose edge returns 403 to it is blocked in every way that matters, and only a real request reveals that.
- Does blocking Google-Extended remove me from Google?
- No. Google-Extended is a robots.txt control token with no crawler behind it. Google crawls with Googlebot and consults the Google-Extended group only to decide whether that content may be used for Gemini training and grounding. Your Search presence and AI Overviews are governed by Googlebot, not by Google-Extended.
- Should I block AI crawlers?
- It depends which kind, and the three kinds do very different things. Blocking training crawlers keeps your content out of future models and changes nothing about today’s answers. Blocking retrieval crawlers removes you from the indexes assistants cite right now. Blocking user-triggered fetchers fails a visitor who explicitly pasted your URL. Most sites that block all three did so with one wildcard rule and meant only the first.
- Is this free?
- Yes. It needs a free Seerly account — no card, and no limit on how many sites you check. Results are cached for a few hours so a site being checked repeatedly is only crawled once, and nothing is stored beyond that.
Look up a specific crawler
Every crawler we check has a reference page: what it does with your content, the exact robots.txt rules to allow or block it, and what blocking it actually costs.
Browse the AI Crawler Directory