AI bot rules are explicit and intentional

Should I allow or block AI crawlers?

Either is defensible — what is not defensible is not knowing. AI crawlers do several different jobs, and treating them as one category means whatever you decide about training data also silently decides whether you appear in AI answers at all.

What it actually is

The bots arriving at your site fall into roughly three groups: those gathering material to train models, those building an index so an assistant can find you later, and those fetching a page right now because a person asked a question about it. They are operated by the same companies and are often assumed to be interchangeable. They are not, and the consequences of refusing each are completely different.

Why it matters

Refusing training crawlers is a legitimate commercial position with a real argument behind it. Refusing the crawler that fetches your page because somebody asked an assistant about your business is simply refusing a customer. The two decisions get conflated constantly, and the second one is usually made by accident while making the first — the traffic disappears and nothing reports the loss.

What good looks like

A named position for each significant AI crawler that reflects what you actually decided, rather than a single sweeping rule applied to everything with "bot" in its name. Anything you have chosen to exclude is excluded deliberately, and anything you want quoting you can reach the pages you want quoted.

Being honest about it

We have no view on which way you should decide and we are not going to pretend there is a correct answer — publishers with valuable archives and businesses wanting to be found in assistants are in genuinely different positions. What we do check is coherence. The failure we care about is a site that meant to protect its content and has instead made itself invisible, or one that never made a decision and does not realise a default was chosen for it.

Where most sites go wrong

The most common failure is the blanket rule adopted during the first wave of concern about AI training, which also excludes the crawlers that would have brought referrals. The second is drift: rules added by different people over two years that now contradict each other, with the effective outcome decided by ordering rather than by anyone's intent. Both look deliberate from the outside and neither is.

← See all 64 checks