Automated visitors can reach your pages

Is my site blocking AI crawlers by accident?

Very possibly, and you would have no way of knowing. Bot protection is usually configured to stop abuse, not to make editorial decisions about answer engines — but the crawlers that would cite you look a lot like the traffic it exists to stop.

What it actually is

Between a visitor and your site there is often a protective layer deciding, in milliseconds, whether each request looks legitimate. It weighs signals no human would notice: the shape of the connection, the reputation of the network, whether the client behaves like a browser. Honest crawlers identify themselves plainly and behave nothing like browsers, which is exactly the profile that layer is tuned to distrust.

Why it matters

An answer engine can only cite a page it managed to load. When your pages are refused, you do not appear as a poor answer — you do not appear at all, and no report anywhere records the absence. This is the failure mode we see most often on otherwise excellent sites, because the protection was switched on deliberately by someone solving a different problem entirely.

What good looks like

Honest automated clients that identify themselves get your actual pages, with the same content a person would see. Protection still applies to traffic that is genuinely abusive, and the crawlers you have deliberately excluded stay excluded — the difference being that every exclusion is one somebody chose.

Being honest about it

One important caveat, because this check is easy to over-read. Most bot-protection vendors permit crawlers that can prove who they are while refusing ones that cannot, so a refusal in our scan is not by itself proof that AI crawlers are blocked — it may simply mean we could not prove our own identity to that vendor. We report what we observed rather than inferring more than we saw. It is a strong signal worth investigating, not a verdict.

Where most sites go wrong

The most common failure is a security setting raised during an attack and never lowered afterwards, which is entirely reasonable at the time and permanent in practice. The second is geographic: rules that refuse whole regions, adopted to reduce nuisance traffic, which also refuse the data centres AI crawlers operate from. In both cases the site works perfectly for every human who visits, which is precisely why nobody finds it.

← See all 64 checks