GPTBot: what it is, what it does with your content, and what blocking it costs

OpenAIModel training

What GPTBot is

GPTBot collects pages to train future models. It does not decide whether OpenAI can cite you in an answer today — that is a different crawler with a different token. Blocking GPTBot is a decision about your content being learned from, not about your visibility.

Operator: OpenAI
robots.txt token: GPTBot
User-agent on the wire: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot
Category: Model training

What blocking it actually costs

Your content is excluded from training future OpenAI models. Does not affect whether ChatGPT can cite you today.

How to block GPTBot

Add this to your robots.txt:

User-agent: GPTBot
Disallow: /

The token must appear on its own User-agent: line. A named group replaces the User-agent: * group rather than adding to it, so anything you also want disallowed for this crawler has to be repeated inside its group.

How to allow GPTBot

User-agent: GPTBot
Allow: /

An explicit allow is worth writing even when you have no blanket block: it documents the decision, and it survives someone later adding a restrictive User-agent: * rule without thinking about AI crawlers.

robots.txt is not the only thing that can block it

A permissive robots.txt does not mean GPTBot can reach you. WAF rules, Cloudflare's bot-management settings, rate limits and country blocks all sit in front of robots.txt and answer first. A site whose robots.txt welcomes GPTBot and whose edge returns 403 to it is blocked in every way that matters — and nothing in robots.txt will tell you so.

This is the gap the AI Crawler Access Checker was built to close: it reads robots.txt and sends a real request carrying the crawler's user-agent, so you see what the crawler sees.

Official documentation

OpenAI documents GPTBot at https://platform.openai.com/docs/bots.

Common questions

What user-agent does GPTBot send?

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot

How do I block GPTBot?

Add a group to robots.txt naming the token exactly:

User-agent: GPTBot
Disallow: /

The group must name GPTBot on its own line. A User-agent: * group does not combine with a named one — under RFC 9309 the most specific matching group replaces the wildcard entirely, it does not inherit from it.

What does blocking GPTBot cost me?

Your content is excluded from training future OpenAI models. Does not affect whether ChatGPT can cite you today.

How do I verify a request really came from GPTBot?

OpenAI publishes the IP ranges it crawls from at https://openai.com/gptbot.json. A user-agent string is trivially spoofed, so anything acting on GPTBot traffic — rate limits, firewall rules, analytics segments — should check the address against that list rather than trusting the header.

Is GPTBot blocked on your site right now?

Reading your own robots.txt only answers half of it — a firewall rule can block GPTBot while robots.txt says it is welcome. The checker tests both.

Check your site free

Related crawlers