AI crawlers

Which AI crawlers visit websites, and what do they do?

The 24 AI crawlers a Bot Appétit scan looks for in your robots.txt: who runs each one, what it's for, and whether its maker says it follows robots.txt. Every fact comes from the maker's own page, and each row says when we last read it, because these change.

Three reasons a crawler visits

Training

It gathers pages to train an AI model. What it reads may shape what the model knows later, with no link back to you.

Search index

It builds the index an AI search product answers from, the way a search engine's crawler always has.

User-triggered

It fetches one page because a person asked an assistant about it, then and there.

The difference decides what a robots.txt rule does. Turn a training crawler away and you ask that company to leave your pages out of what it trains on next. Turn a search crawler away and that product's answers stop drawing on your pages. And some makers say their user-triggered fetchers may not ask robots.txt at all, because a person, not the crawler, chose the page.

The directory

24 crawlers and tokens
AI crawlers
Token Who runs it What it's for Follows robots.txt Last verified
GPTBot OpenAI Their docs about GPTBot (opens in a new tab) Training Yes, says its maker
OAI-SearchBot OpenAI Their docs about OAI-SearchBot (opens in a new tab) Search index Yes, says its maker
ChatGPT-User OpenAI Their docs about ChatGPT-User (opens in a new tab) User-triggered Not always: its maker says robots.txt may not apply to a fetch a person asked for
ClaudeBot Anthropic Their docs about ClaudeBot (opens in a new tab) Training Yes, says its maker
anthropic-ai Anthropic No page from its maker Not documented An old token its maker no longer lists; sites still name it
Claude-User Anthropic Their docs about Claude-User (opens in a new tab) User-triggered Yes, says its maker
Claude-SearchBot Anthropic Their docs about Claude-SearchBot (opens in a new tab) Search index Yes, says its maker
Google-Extended Google Their docs about Google-Extended (opens in a new tab) Training It is a robots.txt rule, not a crawler: another of its maker’s crawlers does the fetching
Google-Agent Google Their docs about Google-Agent (opens in a new tab) User-triggered Not always: its maker says robots.txt may not apply to a fetch a person asked for
PerplexityBot Perplexity Their docs about PerplexityBot (opens in a new tab) Search index Yes, says its maker
Perplexity-User Perplexity Their docs about Perplexity-User (opens in a new tab) User-triggered Not always: its maker says robots.txt may not apply to a fetch a person asked for
CCBot Common Crawl Their docs about CCBot (opens in a new tab) Training Yes, says its maker
Bytespider ByteDance No page from its maker Not documented Not documented: its maker publishes no page about it
Amazonbot Amazon Their docs about Amazonbot (opens in a new tab) Training Yes, says its maker
Amzn-User Amazon Their docs about Amzn-User (opens in a new tab) User-triggered Not always: its maker says robots.txt may not apply to a fetch a person asked for
Amzn-SearchBot Amazon Their docs about Amzn-SearchBot (opens in a new tab) Search index Yes, says its maker
Applebot-Extended Apple Their docs about Applebot-Extended (opens in a new tab) Training It is a robots.txt rule, not a crawler: another of its maker’s crawlers does the fetching
Meta-ExternalAgent Meta Their docs about Meta-ExternalAgent (opens in a new tab) Training Yes, says its maker
Meta-ExternalFetcher Meta Their docs about Meta-ExternalFetcher (opens in a new tab) User-triggered Not always: its maker says robots.txt may not apply to a fetch a person asked for
meta-webindexer Meta Their docs about meta-webindexer (opens in a new tab) Search index Its maker’s page doesn’t say
MistralAI-User Mistral AI Their docs about MistralAI-User (opens in a new tab) User-triggered Yes, says its maker
MistralAI-Index Mistral AI Their docs about MistralAI-Index (opens in a new tab) Search index Its maker’s page doesn’t say
AI2Bot Allen Institute for AI Their docs about AI2Bot (opens in a new tab) Training Its maker’s page doesn’t say
DuckAssistBot DuckDuckGo Their docs about DuckAssistBot (opens in a new tab) Search index Yes, says its maker

"Follows robots.txt" is what each maker's own documentation says, read on 4 October 2026, not something we've measured. A page scan reads whether your robots.txt lets these crawlers in; it doesn't track their visits.

Are they getting in?

A free scan reads your robots.txt and tells you which of these crawlers it turns away, in about 20 seconds — no signup. Is your site blocking ChatGPT, Claude and Perplexity?

Scan your site free