Which AI crawlers visit websites, and what do they do?
The 24 AI crawlers a Bot Appétit scan looks for in your robots.txt: who runs each one, what it's for, and whether its maker says it follows robots.txt. Every fact comes from the maker's own page, and each row says when we last read it, because these change.
Three reasons a crawler visits
Training
Search index
User-triggered
The difference decides what a robots.txt rule does. Turn a training crawler away and you ask that company to leave your pages out of what it trains on next. Turn a search crawler away and that product's answers stop drawing on your pages. And some makers say their user-triggered fetchers may not ask robots.txt at all, because a person, not the crawler, chose the page.
The directory
24 crawlers and tokensShowing all 24
| Token | Who runs it | What it's for | Follows robots.txt | Last verified |
|---|---|---|---|---|
| GPTBot | OpenAI Their docs about GPTBot (opens in a new tab) | Training | Yes, says its maker | |
| OAI-SearchBot | OpenAI Their docs about OAI-SearchBot (opens in a new tab) | Search index | Yes, says its maker | |
| ChatGPT-User | OpenAI Their docs about ChatGPT-User (opens in a new tab) | User-triggered | Not always: its maker says robots.txt may not apply to a fetch a person asked for | |
| ClaudeBot | Anthropic Their docs about ClaudeBot (opens in a new tab) | Training | Yes, says its maker | |
| anthropic-ai | Anthropic No page from its maker | Not documented | An old token its maker no longer lists; sites still name it | |
| Claude-User | Anthropic Their docs about Claude-User (opens in a new tab) | User-triggered | Yes, says its maker | |
| Claude-SearchBot | Anthropic Their docs about Claude-SearchBot (opens in a new tab) | Search index | Yes, says its maker | |
| Google-Extended | Google Their docs about Google-Extended (opens in a new tab) | Training | It is a robots.txt rule, not a crawler: another of its maker’s crawlers does the fetching | |
| Google-Agent | Google Their docs about Google-Agent (opens in a new tab) | User-triggered | Not always: its maker says robots.txt may not apply to a fetch a person asked for | |
| PerplexityBot | Perplexity Their docs about PerplexityBot (opens in a new tab) | Search index | Yes, says its maker | |
| Perplexity-User | Perplexity Their docs about Perplexity-User (opens in a new tab) | User-triggered | Not always: its maker says robots.txt may not apply to a fetch a person asked for | |
| CCBot | Common Crawl Their docs about CCBot (opens in a new tab) | Training | Yes, says its maker | |
| Bytespider | ByteDance No page from its maker | Not documented | Not documented: its maker publishes no page about it | |
| Amazonbot | Amazon Their docs about Amazonbot (opens in a new tab) | Training | Yes, says its maker | |
| Amzn-User | Amazon Their docs about Amzn-User (opens in a new tab) | User-triggered | Not always: its maker says robots.txt may not apply to a fetch a person asked for | |
| Amzn-SearchBot | Amazon Their docs about Amzn-SearchBot (opens in a new tab) | Search index | Yes, says its maker | |
| Applebot-Extended | Apple Their docs about Applebot-Extended (opens in a new tab) | Training | It is a robots.txt rule, not a crawler: another of its maker’s crawlers does the fetching | |
| Meta-ExternalAgent | Meta Their docs about Meta-ExternalAgent (opens in a new tab) | Training | Yes, says its maker | |
| Meta-ExternalFetcher | Meta Their docs about Meta-ExternalFetcher (opens in a new tab) | User-triggered | Not always: its maker says robots.txt may not apply to a fetch a person asked for | |
| meta-webindexer | Meta Their docs about meta-webindexer (opens in a new tab) | Search index | Its maker’s page doesn’t say | |
| MistralAI-User | Mistral AI Their docs about MistralAI-User (opens in a new tab) | User-triggered | Yes, says its maker | |
| MistralAI-Index | Mistral AI Their docs about MistralAI-Index (opens in a new tab) | Search index | Its maker’s page doesn’t say | |
| AI2Bot | Allen Institute for AI Their docs about AI2Bot (opens in a new tab) | Training | Its maker’s page doesn’t say | |
| DuckAssistBot | DuckDuckGo Their docs about DuckAssistBot (opens in a new tab) | Search index | Yes, says its maker |
No crawlers match both filters
"Follows robots.txt" is what each maker's own documentation says, read on 4 October 2026, not something we've measured. A page scan reads whether your robots.txt lets these crawlers in; it doesn't track their visits.
Are they getting in?
A free scan reads your robots.txt and tells you which of these crawlers it turns away, in about 20 seconds — no signup. Is your site blocking ChatGPT, Claude and Perplexity?