Googlebot: what it is, what it does with your content, and what blocking it costs
What Googlebot is
Googlebot builds the retrieval index that Google answers from. This is the category that decides whether you can be cited right now: when someone asks a question and the assistant goes looking for sources, it searches an index this crawler built. Blocking it removes you from those answers.
Operator: Google
robots.txt token: Googlebot
User-agent on the wire: not published by the operator
Category: Live retrieval / search index
What blocking it actually costs
Included because AI Overviews are built on the ordinary Googlebot index: blocking Googlebot removes you from AI Overviews as well as from Search. Probing it is out of scope for this tool — use Search Console.
How to block Googlebot
Add this to your robots.txt:
User-agent: Googlebot
Disallow: /
The token must appear on its own User-agent: line. A named group replaces the User-agent: * group rather than adding to it, so anything you also want disallowed for this crawler has to be repeated inside its group.
How to allow Googlebot
User-agent: Googlebot
Allow: /
An explicit allow is worth writing even when you have no blanket block: it documents the decision, and it survives someone later adding a restrictive User-agent: * rule without thinking about AI crawlers.
robots.txt is not the only thing that can block it
A permissive robots.txt does not mean Googlebot can reach you. WAF rules, Cloudflare's bot-management settings, rate limits and country blocks all sit in front of robots.txt and answer first. A site whose robots.txt welcomes Googlebot and whose edge returns 403 to it is blocked in every way that matters — and nothing in robots.txt will tell you so.
This is the gap the AI Crawler Access Checker was built to close: it reads robots.txt and sends a real request carrying the crawler's user-agent, so you see what the crawler sees.
Official documentation
Google documents Googlebot at https://developers.google.com/search/docs/crawling-indexing/googlebot.
Common questions
- What user-agent does Googlebot send?
Google does not publish an exact user-agent string for Googlebot. Match on the
Googlebottoken in robots.txt rather than on a full string you found elsewhere.- How do I block Googlebot?
Add a group to robots.txt naming the token exactly:
User-agent: Googlebot Disallow: /The group must name
Googleboton its own line. AUser-agent: *group does not combine with a named one — under RFC 9309 the most specific matching group replaces the wildcard entirely, it does not inherit from it.- What does blocking Googlebot cost me?
Included because AI Overviews are built on the ordinary Googlebot index: blocking Googlebot removes you from AI Overviews as well as from Search. Probing it is out of scope for this tool — use Search Console.
- How do I verify a request really came from Googlebot?
Google publishes the IP ranges it crawls from at https://developers.google.com/static/search/apis/ipranges/googlebot.json. A user-agent string is trivially spoofed, so anything acting on Googlebot traffic — rate limits, firewall rules, analytics segments — should check the address against that list rather than trusting the header.
Is Googlebot blocked on your site right now?
Reading your own robots.txt only answers half of it — a firewall rule can block Googlebot while robots.txt says it is welcome. The checker tests both.
Check your site free