AI crawlers aren't blocked

Is your site blocking ChatGPT, Claude and Perplexity?

Quite possibly, and probably without anyone deciding to. If your robots.txt blocks crawlers like GPTBot, ClaudeBot or PerplexityBot, those assistants cannot read your content — so when someone asks them about your industry, you are not in the answer.

What it actually is

Every AI assistant that reads the live web does so with a named crawler, and every one of them checks your robots.txt first. That file is a guest list. The awkward part is that most site owners never wrote theirs: it came with the platform, or an SEO plugin added a blanket rule years ago, or somebody pasted in a "block bad bots" snippet from a forum. The list is enforced regardless of whether anyone meant it.

Why it matters

This is the check with the most direct commercial consequence of any we run. Search rankings are a competition you can lose slowly. Crawler access is binary: if an assistant cannot read the page, you are not a candidate for the answer at all, no matter how good your content is. As more buying research starts inside ChatGPT and Perplexity rather than Google, a blanket block is the difference between being considered and being invisible.

What good looks like

A robots.txt that reaches a deliberate decision about AI crawlers rather than inheriting one. For most businesses that means the major answer-engine crawlers are welcome to your public content, and your genuinely private areas are the only things closed off. The important word is deliberate — a site that has decided to keep AI out for good reasons passes this on intent; a site that is blocking by accident is losing something it never chose to give up.

Being honest about it

Blocking AI crawlers is a legitimate choice, and we are not going to pretend otherwise. Publishers protecting paid archives, sites whose content is the product, anyone who has decided their words should not train a model — those are real positions and we do not mark them as failure for its own sake. What we flag is the gap between what your file says and what you think it says. Most blocks we find were nobody's decision.

Where most sites go wrong

The usual story is inheritance. A blanket rule blocks everything that is not a search engine, written before these crawlers existed and never revisited. Close behind: a rule that blocks one crawler by name while a newer one from the same company walks straight through, because the list was written to a snapshot of the internet that has since moved on. The third pattern is the most frustrating — the file is fine, but a firewall or bot-protection layer in front of the site refuses the crawler anyway, so the welcome mat is out and the door is still locked.

← See all 64 checks