The AI Crawler Directory

36 crawlers and robots.txt control tokens from 21 companies. Grouped by who runs them, because that is how the decision gets made — and tagged by what each one does, because blocking a training crawler and blocking a retrieval crawler have very different consequences.

OpenAI

3 crawlers

Runs three separate crawlers with three different jobs. One wildcard rule blocks all of them, which is almost never what anyone means.

Anthropic

4 crawlers

Separates training, search indexing and user-triggered fetches into distinct tokens.

Google

3 crawlers

Google-Extended is a control token with no crawler behind it — blocking it does not remove you from Search or AI Overviews.

Perplexity

2 crawlers

Microsoft

1 crawler

Meta

3 crawlers

Apple

2 crawlers

Applebot-Extended governs Apple Intelligence training only; Applebot keeps crawling regardless.

Amazon

1 crawler

ByteDance

1 crawler

Widely reported to crawl aggressively and honour robots.txt inconsistently.

Common Crawl

1 crawler

Not an AI company, but its archive is an input to a large number of models.

Mistral AI

1 crawler

DuckDuckGo

1 crawler

Cohere

2 crawlers

You.com

1 crawler

Allen Institute for AI

2 crawlers

Huawei

2 crawlers

Diffbot

1 crawler

ImageSift

1 crawler

Timpi

1 crawler

Webz.io

2 crawlers

Zyte (open source)

1 crawler