Our bot

BotAppetitBot

BotAppetitBot is the scanner behind Bot Appétit, a free tool by Milk Moon Studio that checks whether a page can be read by AI answer engines and search crawlers.

By Milk Moon Studio Last updated

Two kinds of scan, neither a roaming crawler

BotAppetitBot doesn't roam the web, keeps no index, and never goes looking for sites on its own. It fetches only when a person asks.

A page scan reads one page, when someone types its address into Bot Appétit and presses scan, and then it stops. A site scan, which needs an account, reads a sample of up to 50 pages on one site, found from that site's sitemap or, where there isn't one, from the links on its homepage. It runs once, when its owner starts it, and it paces itself: one page at a time, at least a second apart, slowing down when the site pushes back and stopping after repeated refusals. It only picks pages from the site it was asked to scan.

User agents

We send two, and they do different jobs. The scanner makes almost every request: it reads your page, for a page scan and for every page of a site scan alike. The renderer opens the page in a real browser, so you can allow one and refuse the other.

BotAppetitRenderer

The renderer, via Cloudflare Browser Rendering
Signed by Cloudflare
User agent
BotAppetitRenderer/1.0 (+https://scan.milkmoonstudio.com/bot/; on-demand; renders via Cloudflare Browser Rendering)
Key directory
Cloudflare’s, not ours: its signature verifies as Cloudflare.
Standard
RFC 9421 HTTP Message Signatures (Web Bot Auth)

When the renderer runs. Some sites only show their content once JavaScript has run. AI answer engines do not run JavaScript, so to tell you whether your content reaches them we sometimes have to open the page in a real browser and compare. It runs only when the first fetch finds too little text on the page to stand on its own, and never more than once per page. Unlike the scanner, it is a real browser page load, so it also requests the stylesheets and scripts your page references. We block images, video and fonts, which we do not need.

How to verify a request really is us

Anyone can copy a user-agent string, so we sign our requests. Every request the scanner makes to your site carries an RFC 9421 HTTP Message Signature (Web Bot Auth): an Ed25519 signature tagged web-bot-auth, covering the request authority, with our public key published at https://scan.milkmoonstudio.com/.well-known/http-message-signatures-directory.

That directory is itself signed, so nobody can mirror it and claim to be us. If a request's signature verifies against that key, it is genuinely our scanner. If it does not, it is not our scanner — whatever the user-agent says.

The renderer is signed too, but not by us. It runs on Cloudflare Browser Rendering, and Cloudflare attaches its own Web Bot Auth signature to every request that leaves it. It does not carry a signature of ours yet: we have asked Cloudflare to register the renderer as its own agent, and until that is approved we would rather say so than let you read the paragraph above as covering both. So a rendered request carries our user agent above, Cloudflare's signature, and a signature-agent pointing at Cloudflare's key directory rather than ours. It is still verifiable — just as Cloudflare, not as us.

Exactly what one page scan fetches

For a page scan of one page, in one burst, we request:

  • the page itself (one retry if the connection fails outright), and the http:// version of its address, to see whether it redirects to https://
  • /robots.txt, /sitemap.xml, /llms.txt, /.well-known/security.txt, /favicon.ico
  • about a dozen agent-discovery files, most of them under /.well-known/
  • one deliberately non-existent path, to see whether missing pages return a real 404
  • the www/apex variant of the host, once, to check canonicalisation
  • the page again, asking for markdown, and if that does not come back as markdown, the .md version of its address
  • the page's social preview image, if it names one — we read the headers and discard the image
  • up to 12 links from the page, to check they resolve — we read the status code and discard the body unread, and ask once more where a link answers with an error

That is between about 22 and 50 requests, once, for a scan a person explicitly asked for — most of them small files, most of them to your site (the links can lead elsewhere). We do not repeat, poll, or return on a schedule. If the same person scans the same page again, that is another single burst.

A site scan makes that same page scan on each page it reads, one page at a time. A few lookups never reach your server at all: DNS records are read through Cloudflare's public resolver, and Core Web Vitals are measured by Google PageSpeed Insights, so those requests come from Google, not from us.

On some pages there will be one more. When the first fetch suggests the page needs JavaScript to show its content, we open it once in a real browser as BotAppetitRenderer and compare the two. That request behaves like a visitor rather than like the list above: it also loads the stylesheets and scripts your page asks for, so it is the heavier half of what we do. We block images, video and fonts, and we do not scroll, click or submit anything. It happens once, on the same page, in the same burst.

What we keep

The scanner sometimes keeps a copy of a page it scanned: the HTML and the small files a search engine reads, such as robots.txt. It keeps one copy per site for each different set of results, so a rescan that finds the same things replaces the copy rather than adding another, and the pages of a site scan are not copied at all.

Those copies have one job: letting us re-run our own checks against a page we have already seen, so we can tell our scanner changing from your website changing. They are never published, shared or made searchable, and never used to train anything. Ask, and we delete every copy we hold for your site — see Copies of scanned pages.

The renderer keeps nothing. It compares the page it opens with the first fetch, notes how much text JavaScript added, and discards the page.

robots.txt

We read your robots.txt and report on what it says — it is one of the things we check for you. A scan is a fetch a person asked for, so we treat it as a user-directed request rather than crawling: a Disallow rule does not stop the scanner reading the page it was asked to read, and a site scan reads robots.txt for the sitemaps it lists, not to skip pages. When a page a site scan read is one your robots.txt disallows for BotAppetitBot, the report says so and shows the rule that matched — the page is still scanned and still counts. If you would rather we did not fetch your site at all, refuse us at your edge (below) and we will stop.

The renderer holds itself to a stricter rule than that. If your robots.txt disallows us, we do not render — even though the request was user-directed and even though the scanner itself would still fetch the page. The scanner is one request someone asked for; the render additionally pulls in everything your page references, and a site that has told us to go away should not receive the heavy half. Most user-triggered tools, including ones much larger than us, do not draw that line. We would rather draw it.

How to let us in

If a scan came back saying we were refused, this is what opens the door. It only lets us in: your site's AI readiness is exactly as it was, and every finding in the report still stands.

In robots.txt

robots.txt — let BotAppetitBot in
User-agent: BotAppetitBot
Allow: /

A named group beats the wildcard one, so this works even if you disallow everything else. But if you were refused with a 403 or handed a challenge page, the block happened at your edge — before robots.txt was ever read — and the rule below is the one that matters.

At your edge

Every request the scanner makes is signed (see above), so your edge can verify us rather than trust a string. Failing that, our user agent is stable and contains BotAppetitBot.

  • Cloudflare, Pro and up: a WAF custom rule with the Skip action. Custom rules run before Super Bot Fight Mode, which is what makes the exception stick.
  • Cloudflare, Free: if Bot Fight Mode is on, there is no exception to make — Cloudflare's own documentation says it "cannot be bypassed or skipped using WAF custom rules or Page Rules", because it runs outside the rules engine. Turning it off, or moving to Super Bot Fight Mode, are the only two answers.
  • Akamai, and other bot managers: allow the user agent above, or our web-bot-auth signature where the product supports it.

We do not claim your vendor has verified us, and you should not have to take our word for anything: the signature above is checkable, and that is the point of publishing it.

Once you've let us in, go back to your report and rescan the page.

Not sure what's blocking us, or how to change it? We can look. No pitch, just a plain-English chat.

Book a call

How to block us

Refuse us at your edge: answer our user agent or our signature with a 403. A page scan then reports the refusal and scores nothing, and a site scan slows down and stops after repeated refusals. We would rather you did not — but a block is a legitimate answer, and we will not work around one.

robots.txt stops the renderer, not the scanner (see above). A rule aimed at the scanner binds the renderer too:

robots.txt — block BotAppetitBot
User-agent: BotAppetitBot
Disallow: /

If you are happy for us to read your page but not to open it in a browser, refuse only the renderer. Everything else in your report is unaffected; the one check that needs a browser will simply say we could not look.

robots.txt — block BotAppetitRenderer
User-agent: BotAppetitRenderer
Disallow: /

Contact

Something wrong, or want us to stop? jakes@milkmoonstudio.com — a real person reads it.