The AI Crawler Directory
36 crawlers and robots.txt control tokens from 21 companies. Grouped by who runs them, because that is how the decision gets made — and tagged by what each one does, because blocking a training crawler and blocking a retrieval crawler have very different consequences.
OpenAI
3 crawlersRuns three separate crawlers with three different jobs. One wildcard rule blocks all of them, which is almost never what anyone means.
Anthropic
4 crawlersSeparates training, search indexing and user-triggered fetches into distinct tokens.
Google-Extended is a control token with no crawler behind it — blocking it does not remove you from Search or AI Overviews.
Perplexity
2 crawlersMicrosoft
1 crawlerMeta
3 crawlersApple
2 crawlersApplebot-Extended governs Apple Intelligence training only; Applebot keeps crawling regardless.
Amazon
1 crawlerByteDance
1 crawlerWidely reported to crawl aggressively and honour robots.txt inconsistently.
Common Crawl
1 crawlerNot an AI company, but its archive is an input to a large number of models.