Two kinds of scan, neither a roaming crawler
BotAppetitBot doesn't roam the web, keeps no index, and never goes looking for sites on its own. It fetches only when a person asks.
A page scan reads one page, when someone types its address into Bot Appétit and presses scan, and then it stops. A site scan, which needs an account, reads a sample of up to 50 pages on one site, found from that site's sitemap or, where there isn't one, from the links on its homepage. It runs once, when its owner starts it, and it paces itself: one page at a time, at least a second apart, slowing down when the site pushes back and stopping after repeated refusals. It only picks pages from the site it was asked to scan.
User agents
We send two, and they do different jobs. The scanner makes almost every request: it reads your page, for a page scan and for every page of a site scan alike. The renderer opens the page in a real browser, so you can allow one and refuse the other.
BotAppetitBot
The scanner- User agent
-
BotAppetitBot/1.0 (+https://scan.milkmoonstudio.com/bot/; on-demand) - Key directory
- https://scan.milkmoonstudio.com/.well-known/http-message-signatures-directory
- Standard
- RFC 9421 HTTP Message Signatures (Web Bot Auth)
BotAppetitRenderer
The renderer, via Cloudflare Browser Rendering- User agent
-
BotAppetitRenderer/1.0 (+https://scan.milkmoonstudio.com/bot/; on-demand; renders via Cloudflare Browser Rendering) - Key directory
- Cloudflare’s, not ours: its signature verifies as Cloudflare.
- Standard
- RFC 9421 HTTP Message Signatures (Web Bot Auth)
When the renderer runs. Some sites only show their content once JavaScript has run. AI answer engines do not run JavaScript, so to tell you whether your content reaches them we sometimes have to open the page in a real browser and compare. It runs only when the first fetch finds too little text on the page to stand on its own, and never more than once per page. Unlike the scanner, it is a real browser page load, so it also requests the stylesheets and scripts your page references. We block images, video and fonts, which we do not need.
How to verify a request really is us
Anyone can copy a user-agent string, so we sign our requests. Every request the
scanner makes to your site carries an
RFC 9421 HTTP Message Signature (Web Bot
Auth): an Ed25519 signature tagged web-bot-auth, covering the request
authority, with our public key published at
https://scan.milkmoonstudio.com/.well-known/http-message-signatures-directory.
That directory is itself signed, so nobody can mirror it and claim to be us. If a request's signature verifies against that key, it is genuinely our scanner. If it does not, it is not our scanner — whatever the user-agent says.
The renderer is signed too, but not by us. It runs on Cloudflare Browser Rendering, and
Cloudflare attaches its own Web Bot Auth signature to every request that leaves it. It does not carry a
signature of ours yet: we have asked Cloudflare to register the renderer as its own agent, and until that
is approved we would rather say so than let you read the paragraph above as covering both. So a rendered
request carries our user agent above, Cloudflare's signature, and a signature-agent pointing
at Cloudflare's key directory rather than ours. It is still verifiable — just as Cloudflare, not as us.
Exactly what one page scan fetches
For a page scan of one page, in one burst, we request:
-
the page itself (one retry if the connection fails outright), and the
http://version of its address, to see whether it redirects tohttps:// -
/robots.txt,/sitemap.xml,/llms.txt,/.well-known/security.txt,/favicon.ico - about a dozen agent-discovery files, most of them under
/.well-known/ - one deliberately non-existent path, to see whether missing pages return a real 404
- the
www/apex variant of the host, once, to check canonicalisation -
the page again, asking for markdown, and if that does not come back as markdown, the
.mdversion of its address - the page's social preview image, if it names one — we read the headers and discard the image
- up to 12 links from the page, to check they resolve — we read the status code and discard the body unread, and ask once more where a link answers with an error
That is between about 22 and 50 requests, once, for a scan a person explicitly asked for — most of them small files, most of them to your site (the links can lead elsewhere). We do not repeat, poll, or return on a schedule. If the same person scans the same page again, that is another single burst.
A site scan makes that same page scan on each page it reads, one page at a time. A few lookups never reach your server at all: DNS records are read through Cloudflare's public resolver, and Core Web Vitals are measured by Google PageSpeed Insights, so those requests come from Google, not from us.
On some pages there will be one more. When the first fetch suggests the page needs
JavaScript to show its content, we open it once in a real browser as BotAppetitRenderer and
compare the two. That request behaves like a visitor rather than like the list above: it also loads the
stylesheets and scripts your page asks for, so it is the heavier half of what we do. We block images,
video and fonts, and we do not scroll, click or submit anything. It happens once, on the same page, in
the same burst.
What we keep
The scanner sometimes keeps a copy of a page it scanned: the HTML and the small files a
search engine reads, such as robots.txt. It keeps one copy per site for each different set
of results, so a rescan that finds the same things replaces the copy rather than adding another, and the
pages of a site scan are not copied at all.
Those copies have one job: letting us re-run our own checks against a page we have already seen, so we can tell our scanner changing from your website changing. They are never published, shared or made searchable, and never used to train anything. Ask, and we delete every copy we hold for your site — see Copies of scanned pages.
The renderer keeps nothing. It compares the page it opens with the first fetch, notes how much text JavaScript added, and discards the page.
robots.txt
We read your robots.txt and report on what it says — it is one of the things we check for
you. A scan is a fetch a person asked for, so we treat it as a user-directed request rather than
crawling: a Disallow rule does not stop the scanner reading the page it was asked to read,
and a site scan reads robots.txt for the sitemaps it lists, not to skip pages. When a page a
site scan read is one your robots.txt disallows for BotAppetitBot, the report says so and
shows the rule that matched — the page is still scanned and still counts. If you would rather we did not
fetch your site at all, refuse us at your edge (below) and we will stop.
The renderer holds itself to a stricter rule than that. If your robots.txt
disallows us, we do not render — even though the request was user-directed and even though the scanner
itself would still fetch the page. The scanner is one request someone asked for; the render additionally
pulls in everything your page references, and a site that has told us to go away should not receive the
heavy half. Most user-triggered tools, including ones much larger than us, do not draw that line. We would
rather draw it.
How to let us in
If a scan came back saying we were refused, this is what opens the door. It only lets us in: your site's AI readiness is exactly as it was, and every finding in the report still stands.
In robots.txt
User-agent: BotAppetitBot Allow: /
A named group beats the wildcard one, so this works even if you disallow everything else. But if you
were refused with a 403 or handed a challenge page, the block happened at your edge — before
robots.txt was ever read — and the rule below is the one that matters.
At your edge
Every request the scanner makes is signed (see above), so your edge can verify us rather than trust a
string. Failing that, our user agent is stable and contains BotAppetitBot.
- Cloudflare, Pro and up: a WAF custom rule with the Skip action. Custom rules run before Super Bot Fight Mode, which is what makes the exception stick.
- Cloudflare, Free: if Bot Fight Mode is on, there is no exception to make — Cloudflare's own documentation says it "cannot be bypassed or skipped using WAF custom rules or Page Rules", because it runs outside the rules engine. Turning it off, or moving to Super Bot Fight Mode, are the only two answers.
- Akamai, and other bot managers: allow the user agent above, or our
web-bot-authsignature where the product supports it.
We do not claim your vendor has verified us, and you should not have to take our word for anything: the signature above is checkable, and that is the point of publishing it.
Once you've let us in, go back to your report and rescan the page.
Not sure what's blocking us, or how to change it? We can look. No pitch, just a plain-English chat.
Book a callHow to block us
Refuse us at your edge: answer our user agent or our signature with a 403. A page scan then
reports the refusal and scores nothing, and a site scan slows down and stops after repeated refusals. We
would rather you did not — but a block is a legitimate answer, and we will not work around one.
robots.txt stops the renderer, not the scanner (see above). A rule aimed at the scanner
binds the renderer too:
User-agent: BotAppetitBot Disallow: /
If you are happy for us to read your page but not to open it in a browser, refuse only the renderer. Everything else in your report is unaffected; the one check that needs a browser will simply say we could not look.
User-agent: BotAppetitRenderer Disallow: /
Contact
Something wrong, or want us to stop? jakes@milkmoonstudio.com — a real person reads it.