The AEO and SEO glossary
What do AEO and AI-readiness terms mean?
Every term you'll meet on Bot Appétit, from AEO to Web Bot Auth, in a sentence or two of plain English. Where a term has its own check, the entry links to the page that explains it in full.
Showing every term, A to Z
#
- 301 redirect map Checklist vocabulary
- A list pairing old addresses with their new ones, so every moved page redirects permanently.
- Checked by: Reached without a redirect chain
- 404 page / soft 404 The page
- A 404 is the status for "this page does not exist". A soft 404 is a missing page that answers with a normal 200 status instead, so crawlers may treat junk as real pages.
- Checked by: Missing pages return a real 404
A
- Advanced Our own words
- The course of 11 agent-readiness checks ("Agent & MCP readiness"). They run on every scan and are shown, but they do not count toward the headline score, because most sites do not need them yet.
- Read more about Advanced
- AEO (answer-engine optimisation) The field
- Making a site easy for AI answer engines to find, read, understand and cite, so a page can be quoted in an AI-written answer. It overlaps with SEO but is judged by whether a machine can consume the page, not by where it ranks.
- Read more about AEO (answer-engine optimisation)
- Agent & MCP readiness Agent readiness
- Our name for the Advanced course: the emerging, mostly draft signals AI agents use to act on a site. Most sites do not need them yet.
- Read more about Agent & MCP readiness
- Agent Skills index Agent readiness
- A file at /.well-known/agent-skills/index.json listing reusable "skills" (instructions and resources) an agent can load to work with a site or service. It comes from a discovery proposal, not a finished standard.
- Checked by: Agent Skills index
- AI agent Agent readiness
- An AI system that acts for a person (searching, booking, buying, filling in forms) rather than only answering. Agents need machine-readable ways to find and use what a site offers.
- AI catalog Agent readiness
- A single index file (/.well-known/ai-catalog.json) listing the AI-facing things a site publishes (MCP servers, agent cards, skills), so an agent finds them in one request. The specification is a draft; an index with nothing in it tells an agent nothing.
- Checked by: AI catalog
- AI crawler Bots and access
- A crawler run by an AI company. Some collect training data, some build an AI search index, and some fetch a page only when a user asks.
- Checked by: AI crawlers aren't blocked Read more about AI crawler
- AI discovery via DNS Agent readiness
- Records in a domain's DNS that point agents at its AI endpoints without fetching a page. The formats come from IETF drafts and informal conventions, and adoption is early.
- Checked by: AI discovery via DNS
- AI Overview / AI Mode The field
- Google Search features that put an AI-generated summary (AI Overview) or a conversational AI answer (AI Mode) above or in place of the usual results, with links out to sources. Google says there are no special extra requirements to appear in them.
- Checked by: AI crawlers aren't blocked
- Alt text The page
- A text description of an image in its alt attribute, read by screen readers and machines that cannot see the picture. An empty alt="" correctly marks a decorative image.
- Checked by: Images have descriptive alt text
- Analytics / GA4 / GTM Checklist vocabulary
- Analytics measures visits. GA4 is Google Analytics 4; GTM (Google Tag Manager) is a container that loads tags such as analytics.
- Checked by: Analytics is installed
- Anchor text The page
- The clickable words of a link. Descriptive anchors tell machines what is on the other end; "click here" tells them nothing.
- Checked by: Descriptive anchor text
- Answer engine The field
- A service that replies to a question with a written answer (often citing sources) instead of a list of links: ChatGPT, Claude, Perplexity, Google's AI features, Microsoft Copilot.
- Checked by: AI crawlers aren't blocked
- Answer-first The page
- Opening a page or section with a direct answer to the question its heading asks, then the detail. It is the AEO writing practice our own check pages follow.
- Checked by: Content is scannable in self-contained chunks
- API Agent readiness
- An interface one program uses to talk to another, as opposed to pages a person reads.
- Checked by: API catalog
- API catalog Agent readiness
- A machine-readable list of a site's APIs at /.well-known/api-catalog (RFC 9727), so agents can find them.
- Checked by: API catalog
- ARIA / ARIA landmarks The page
- Attributes that add accessibility meaning to HTML for assistive technology; ARIA landmark roles are the attribute form of semantic regions.
- Checked by: Semantic HTML landmarks
- auth.md Agent readiness
- A Markdown file at a domain's root (/auth.md) describing how an AI agent can register with a service for a user: flows, scopes and endpoints. An open protocol authored by WorkOS, new in 2026.
- Checked by: Auth.md agent registration
- Author byline / publisher Structured data
- The named person or organisation responsible for a page. It helps a machine only when it is declared in markup, not just shown as text.
- Checked by: Author / publisher is clear
B
- Blocked report / unscored report ("No score for this one") Our own words
- When a site refuses our scanner, we report the site-level signals we could reach and publish no score at all, never a zero.
- Checked by: Automated visitors can reach your pages
- Bot protection / WAF (web application firewall) Bots and access
- Security services that sit in front of a site and refuse traffic they judge automated or hostile. Most let verified search and AI crawlers through while refusing unverified bots, which is why being refused is not proof AI crawlers are blocked.
- Checked by: Automated visitors can reach your pages
- BotAppetitBot / BotAppetitRenderer Bots and access
- The names our scanner and our JavaScript renderer send, so a site can recognise, allow or block each separately.
- Checked by: Automated visitors can reach your pages Read more about BotAppetitBot / BotAppetitRenderer
- Breadcrumb Structured data
- The trail of links at the top of a page showing where it sits in the site. BreadcrumbList is its machine-readable form.
- Checked by: Breadcrumb structured data
- BreadcrumbList Structured data
- The schema.org type that describes a page's position in the site's hierarchy (Home › Section › Page), giving machines the site's structure.
- Checked by: Breadcrumb structured data
- Broken link The page
- A link whose target answers "not found" or a server error.
- Checked by: No broken links
C
- Caching headers (Cache-Control, ETag, Last-Modified, Expires) Performance
- Headers that let browsers and CDNs keep a copy of a page and check whether it changed, instead of downloading it again.
- Checked by: Caching or validation headers are set
- Canonical URL / canonical tag The page
- The address a page declares as its official version (a rel="canonical" link or header), so duplicates of the same content are treated as one page.
- Checked by: This page declares its canonical URL
- Canonicalization (www vs non-www) The page
- A site settling on one host name, either with or without www., so the same content isn't served at two addresses.
- Checked by: The site settles on one host (www vs non-www)
- Category / course Our own words
- A group of related checks: AEO / content, AI access, AI signals, Structured data, SEO, Links, Security (scored), plus Advanced, Good practice and Performance. "Course" is the menu name for a category.
- CDN (content delivery network) Performance
- A network of servers around the world that keeps copies of a site close to visitors.
- Checked by: Caching or validation headers are set
- Changelog ("What we changed") Our own words
- Our public, dated record of changes to how we judge pages, in plain language and without weights.
- Read more about Changelog ("What we changed")
- Check Our own words
- One test the scanner runs, answering one question about whether a machine can consume the page. There are 66 in all.
- Chef's kiss (verdict, top tier) Our own words
- The best verdict. "Nothing to send back — AI engines can read you loud and clear."
- Chunking / passage The page
- Writing a page as self-contained headed sections, each about one idea. Answer engines lift passages, not whole pages, so a section with clear edges is easier to quote.
- Checked by: Content is scannable in self-contained chunks
- Citation (in an AI answer) The field
- A link or named source an AI answer shows for a claim. Being cited is the AEO equivalent of ranking, and it depends on the page being readable and attributable.
- Checked by: AI crawlers aren't blocked
- Cleans the plate (verdict, second tier) Our own words
- "A strong plate. A little polish here and there and this is a sure thing."
- Cloudflare Turnstile Bots and access
- Cloudflare's check that a visitor is a person, usually invisible, that runs before each scan to stop automated abuse of the scanner. A checkbox appears only when Cloudflare insists.
- CLS (Cumulative Layout Shift) Performance
- How much the layout jumps around while the page loads. Google counts 0.1 or less as good.
- Checked by: Images are optimised
- Colour contrast (WCAG AA) Checklist vocabulary
- How distinguishable text is from its background. WCAG AA is the accessibility level most sites aim for.
- Content negotiation Files and signals
- The HTTP mechanism where a client says which formats it prefers (the Accept header) and the server picks the one to send.
- Checked by: Clean content for agents (Markdown)
- Content type / entity type (Article, Product, LocalBusiness, Person…) Structured data
- A schema.org type that commits to what a page is about, beyond the generic WebSite or WebPage every page carries.
- Checked by: Right schema type for the page
- Content-Signal Bots and access
- A line in robots.txt, or a response header, that states how content may be used after it is fetched: search (search results), ai-input (used as input to an AI answer) and ai-train (used to train models). Introduced by Cloudflare in 2025; it expresses a preference and does not block anything.
- Checked by: Content usage preferences declared
- @context / @type Structured data
- The two keys that make a JSON-LD entity mean something: @context says which vocabulary is used (usually schema.org), and @type says what kind of thing it is. Without them an engine cannot tell what a block describes.
- Checked by: Structured data is valid
- Core Web Vitals Performance
- Google's three measures of how a page feels to real users: loading (LCP), responsiveness (INP) and visual stability (CLS). We measure them live and show them in their own Performance card, outside the headline score.
- Checked by: Largest Contentful Paint (LCP) Read more about Core Web Vitals
- Couldn't judge Our own words
- We tried and could not look: the request was refused, timed out or the page didn't load. It stays out of the score rather than being guessed.
- Course by course Our own words
- The report section listing every check result, grouped by category, each expandable to "What we found", "Why it matters" and "What good looks like".
- Coverage Our own words
- How many pages we could actually read and how many checks we could judge in a site scan, stated alongside the score.
- Crawl budget The field
- The number of pages a crawler is willing to fetch from a site in a given time. Dead links and redirects spend it on nothing.
- Checked by: No broken links
- Crawl-delay Bots and access
- An unofficial robots.txt line asking a crawler to wait between requests. It is not part of RFC 9309; some crawlers honour it, and our site scan does, up to a ceiling.
- Crawler / bot / spider Bots and access
- A program that fetches web pages automatically. Search engines and AI companies each run their own, and each announces itself with a user agent name.
- Checked by: AI crawlers aren't blocked
- Crawler vs fetcher vs training bot Bots and access
- Three jobs an AI company's bots do. A training bot collects pages that may be used to train future models. A search bot (crawler) indexes pages so an AI search product can find and cite them. A user fetcher visits a single page because a person just asked the assistant something. Vendors publish separate names for each, and some say robots.txt may not apply to user fetchers because a person started the request.
- Checked by: AI bot rules are explicit and intentional
- Custom domain Checklist vocabulary
- A site's own address, such as example.com, rather than the platform's default subdomain.
- Checked by: The site settles on one host (www vs non-www)
D
- Disallow / Allow Bots and access
- The two rule words in robots.txt. Disallow: / asks a crawler to fetch nothing on the site; a more specific Allow can override a broader Disallow for the same crawler.
- Checked by: robots.txt is present and not over-blocking
- DNS / DNS-over-HTTPS Agent readiness
- DNS is the internet's address book, turning names into server addresses and holding text records. DNS-over-HTTPS is a way to query it over an encrypted web connection, which is how we read those records.
- Checked by: AI discovery via DNS
- Do-first shortlist Checklist vocabulary
- The checklist's short list of the highest-value items to start with.
- Read more about Do-first shortlist
- Duplicate content The page
- The same content reachable at more than one address. Machines have to guess which is the real one.
- Checked by: This page declares its canonical URL
E
- E-E-A-T Structured data
- Google's shorthand for Experience, Expertise, Authoritativeness and Trustworthiness, used in its guidance on helpful content. Google says it is not itself a ranking factor. We use it only to describe clear authorship signals.
- Checked by: Author / publisher is clear
- Entity Structured data
- A single real-world thing that structured data describes, such as a business, a person or a product. Machines try to match the same entity across different sources.
- Checked by: Organization schema with sameAs
F
- FAQPage Structured data
- The schema.org type for a page with a list of questions and the site's own answers. It labels Q&A content so a machine can lift question-and-answer pairs.
- Checked by: FAQ / Q&A schema
- Favicon / touch icon (webclip) Files and signals
- The small icon for a site that browsers show in tabs, and that some search results and AI answer interfaces show next to a link. A touch icon is the larger version phones use for home-screen shortcuts.
- Checked by: A favicon is set
- FCP (First Contentful Paint) Performance
- How long until the first text or image appears. A supporting metric, not a Core Web Vital.
- Checked by: Largest Contentful Paint (LCP)
- Field data / CrUX (Chrome User Experience Report) Performance
- Measurements from real Chrome users visiting the page over the last 28 days, published by Google. A page with little traffic has none.
- Checked by: Largest Contentful Paint (LCP)
G
- Gauge Our own words
- The dial that shows the headline score.
- Generative AI The field
- AI that produces new text, images or code rather than only classifying or ranking. In this product it is shorthand for the systems that read pages to answer or to train.
- GEO (generative engine optimisation) The field
- Another name for optimising content to be surfaced by generative AI search. It comes from a 2023 research paper and is used interchangeably with AEO in the industry.
- Good practice Our own words
- 6 checks worth getting right that don't change whether a machine can find, fetch, parse, trust or cite the page. Measured and shown, never scored.
- Checked by: Title fits a search snippet
- Google Safe Browsing Security
- Google's list of sites known for malware, phishing or unwanted software, used by browsers to show warnings. We look a page up with a privacy-preserving method and leave the check out rather than guess when the answer is unknown.
- Checked by: Not flagged for malware or phishing
- Google Search Console / Bing Webmaster Tools Checklist vocabulary
- Free dashboards from Google and Microsoft where a verified site owner sees index status, errors and search performance. Verification proves you own the site.
- Checked by: Search Console / Bing verification
- @graph Structured data
- A JSON-LD wrapper that holds several related entities in one block, for example the organisation, the website and the page.
- Checked by: Structured data is valid
- Grounding / RAG (retrieval-augmented generation) The field
- An AI system fetching real documents at the moment it answers, and basing its answer on them, instead of relying only on what it learned in training. This is the moment a page gets read and cited.
- Checked by: Content usage preferences declared
H
- Heading / H1 / heading hierarchy The page
- Headings (H1 to H6) are the titles of a page's sections. The H1 names the page's topic, and the levels below should nest in order so a machine can read the outline.
- Checked by: Exactly one H1
- Headless browser / renderer The page
- A real browser run by a program, with no screen, used to load a page with JavaScript the way a person's browser would. Our renderer uses Cloudflare Browser Rendering, and only when the raw HTML looks thin.
- Checked by: Your content survives without JavaScript Read more about Headless browser / renderer
- hreflang The page
- Link annotations that tell engines which language or regional version of a page to show to whom. Each value must be a valid language tag.
- Checked by: hreflang for language variants
- HSTS (HTTP Strict Transport Security) Security
- A header (RFC 6797) that tells browsers to use only HTTPS for a site for a set period, so a visitor cannot be quietly downgraded to plain HTTP.
- Checked by: Core security headers are set
- HTML / raw HTML The page
- HTML is the markup a web page is written in. The raw HTML is what the server sends before any JavaScript runs, and it is all most AI crawlers ever read.
- Checked by: Your answer is in the HTML
- HTTP header / response header Files and signals
- Lines of information a server sends before the page itself: its type, caching rules, security policies, links to related files. Machines read them before any HTML.
- Checked by: Core security headers are set
- HTTP message signatures Bots and access
- A published standard (RFC 9421) for attaching a cryptographic signature to an HTTP request or response, so the receiver can check who sent it and that it was not altered. Web Bot Auth is built on it.
- Checked by: Web Bot Auth (operates a signed bot)
- HTTP status code The page
- The three-digit number a server sends with every response: 200 means OK, 3xx means moved, 404 means not found, 5xx means the server failed. Machines act on the number before reading any content.
- Checked by: Reached without a redirect chain
- HTTP to HTTPS redirect Security
- The insecure http:// address sending visitors on to the https:// one, ideally permanently.
- Checked by: HTTP redirects to HTTPS
- HTTPS / TLS certificate Security
- HTTPS is the encrypted version of HTTP. The certificate proves the site is who it says it is; an expired or invalid one triggers browser warnings.
- Checked by: Served over a secure HTTPS connection
I
- Iframe (embedded frame) The page
- A separate page shown inside another page. A machine reading the outer page does not treat the frame's content as part of it.
- Checked by: Your content survives without JavaScript
- Image formats: WebP / AVIF Performance
- Modern image formats that are usually much smaller than JPEG or PNG at the same quality.
- Checked by: Images are optimised
- Impact / effort (checklist ratings) Checklist vocabulary
- The checklist's own labels for how much an item matters and how much work it is. They are editorial ratings, not Bot Appétit's scoring.
- Read more about Impact / effort (checklist ratings)
- Indexing / indexable The field
- A search engine storing a page so it can be shown in results. An indexable page is one that does not ask to be left out.
- Checked by: Not accidentally set to noindex
- Inline statistics and citations / entity consistency Checklist vocabulary
- Backing claims with specific figures and named sources, and using the same names for the same things throughout. The research paper that named GEO measured both statistics and citations as raising how visible a page was in AI-written answers.
- Checked by: Topic is consistent across title, H1 and description
- INP (Interaction to Next Paint) Performance
- How quickly a page responds visibly after a tap, click or key press, across the visit. Google counts 200 ms or less as good. It replaced First Input Delay as a Core Web Vital in 2024, and it can only be measured from real users.
- Checked by: Interaction to Next Paint (INP)
- Internal link / external link The page
- A link to another page on the same site, or to a different site. Internal links are how crawlers move through a site.
- Checked by: Sensible internal/external link mix
J
- JavaScript rendering / client-side rendering The page
- A page whose content is built in the browser by JavaScript after the HTML arrives. Search engines like Google render it; most AI answer engines read the HTML and never run the script.
- Checked by: Your content survives without JavaScript
- JavaScript shell The page
- A page that sends almost no text in its HTML and relies on scripts to fill it in. To a crawler that doesn't run JavaScript it looks empty.
- Checked by: Your answer is in the HTML
- JSON-LD Structured data
- The format most sites use for structured data: a block of JSON inside the HTML, a W3C Recommendation. Machines read it without having to interpret the visible page.
- Checked by: Structured data (JSON-LD) is present
K
- Keyword / primary keyword / topic consistency The page
- The words people search for, and the main one a page is about. Topic consistency means the title, H1 and description agree on it.
- Checked by: Topic is consistent across title, H1 and description
- Keyword map Checklist vocabulary
- A plan that gives each page one main search term, so pages don't compete with each other.
- Read more about Keyword map
L
- Lab data / lab test Performance
- A single simulated page load on a set device and connection. It is repeatable but is not what real users experienced. Ours is a simulated mobile run.
- Checked by: Largest Contentful Paint (LCP)
- Lab score (0 to 100) Performance
- PageSpeed's own performance score from the lab run, shown on our Performance card. It is Google's number, not ours, and never enters the headline.
- Checked by: Page weight and render-blocking
- lang attribute / charset The page
- The lang attribute declares the page's human language; the charset declares how its characters are encoded (normally UTF-8). Both help machines read and pronounce the text correctly.
- Checked by: Language and charset declared
- Language tag (BCP 47) The page
- The standard codes for languages and regions, such as en, en-GB or pt-BR.
- Checked by: hreflang for language variants
- Lazy-loading Performance
- Waiting to download an image until the visitor scrolls near it. It saves weight below the fold, and it delays an image that is already in view.
- Checked by: Images are optimised
- LCP (Largest Contentful Paint) Performance
- How long until the biggest piece of content in view (usually the hero image or headline) has loaded. Google counts 2.5 seconds or less as good.
- Checked by: Images are optimised
- Let us in Our own words
- The panel on a blocked report explaining, for the site's owner, how our scanner identifies itself so they can allow it, and that allowing it fixes nothing else.
- Checked by: Automated visitors can reach your pages Read more about Let us in
- Lighthouse Performance
- Google's open-source tool that runs a lab test and scores performance, accessibility and SEO. PageSpeed Insights runs it.
- Checked by: Page weight and render-blocking
- Link header Files and signals
- An HTTP header (RFC 8288) that points to related resources, for example the sitemap, the canonical URL, or a describedby document such as llms.txt, so an agent can find them without parsing the HTML.
- Checked by: Resource-discovery Link headers
- LLM (large language model) The field
- The kind of AI model behind ChatGPT, Claude, Gemini and similar tools: trained on large amounts of text to read and write language. An LLM reading your page sees text and markup, not your design.
- Checked by: llms.txt is present and valid
- llms-full.txt Files and signals
- A companion convention: one file holding a site's whole text content in Markdown, so an AI tool can load all of it from one address. It started with Mintlify and Anthropic's docs and is not part of the llmstxt.org proposal. We do not check for it.
- Checked by: llms.txt is present and valid
- llms.txt Files and signals
- A plain-text Markdown file at the root of a site (/llms.txt) that gives AI tools a short, curated map of the site's key content. Proposed by Jeremy Howard in September 2024; an H1 naming the site is its one required element. Adoption is early.
- Checked by: llms.txt is present and valid
M
- Machine-readable / machine readability The field
- Information a program can parse without guessing: text in the HTML, dates in a set format, facts in structured data. The product's single question is whether a machine can consume the page.
- Machine-readable date (datePublished, dateModified, <time datetime>) Structured data
- A published or updated date in a fixed format in structured data, article meta tags or a time element, so a machine does not have to guess from prose.
- Checked by: A machine-readable published / updated date
- Markdown Files and signals
- A plain-text format that marks headings, lists and links with simple symbols. It carries a page's content without layout code, which makes it cheap and clean for an AI model to read.
- Checked by: Clean content for agents (Markdown)
- Markdown negotiation / Markdown for agents Files and signals
- A site answering a request that asks for Markdown (Accept: text/markdown) with a Markdown version of the page instead of HTML, or serving one at the same address with .md on the end. Cloudflare offers it as a zone feature; it is optional and few platforms support it.
- Checked by: Clean content for agents (Markdown)
- MCP (Model Context Protocol) Agent readiness
- An open standard for connecting AI applications to outside data, tools and workflows. A site or service offers them through an MCP server.
- Checked by: MCP Server Card
- MCP server Agent readiness
- A service that exposes tools and data to AI applications over MCP.
- Checked by: MCP Server Card
- MCP Server Card Agent readiness
- A small file at a well-known address that advertises a site's MCP server so agents can discover it. The proposal (SEP-2127) is still a draft.
- Checked by: MCP Server Card
- Measured, not scored Our own words
- The report group for Advanced, Good practice and Performance: real findings about the page that do not move the headline score.
- Meta description The page
- A short summary of the page in the HTML head, often shown as the snippet under a search result. It is not visible on the page itself.
- Checked by: Compelling meta description
- Meta refresh The page
- A redirect written into the page's HTML (<meta http-equiv="refresh">) instead of sent by the server. It is slower, and crawlers may ignore it.
- Checked by: No meta-refresh redirect
- Meta robots / robots meta tag The page
- A tag in a page's HTML head that tells crawlers whether they may index the page or follow its links. Unlike robots.txt, it works per page.
- Checked by: Not accidentally set to noindex
- Microdata Structured data
- An older way to embed schema.org data as attributes on visible HTML elements instead of a separate JSON-LD block. We read it too.
- Checked by: FAQ / Q&A schema
- Minification Performance
- Removing whitespace and comments from code files to make them smaller.
- Checked by: Page weight and render-blocking
- Mixed content Security
- An HTTPS page that loads some of its scripts, styles, frames, images or media over insecure HTTP. Browsers block the risky kinds and try to upgrade the rest.
- Checked by: No mixed (insecure) content
N
- Needs seasoning (verdict, middle tier) Our own words
- "The bones are good, but answer engines are leaving hungry."
- nofollow The page
- A link attribute asking search engines not to pass credit through that link. A page that marks most of its links nofollow looks unusual.
- Checked by: Sensible internal/external link mix
- noindex The page
- The robots directive asking search engines to leave a page out of their results. It is often left on by mistake after launch. We also set it on our own reports, because they hold findings about other people's sites.
- Checked by: Not accidentally set to noindex
- Not applicable (omitted) Our own words
- A check that doesn't apply to this page (an image check on a page with no images) is left out entirely, with no row.
- Not compared / "our scoring changed" Our own words
- Two scores we won't subtract, because the pinned pages changed between them. A change in our scoring is still shown, marked "our scoring changed", with a link to what we changed.
O
- OAuth Agent readiness
- The standard way to let an app act on your behalf without giving it your password. Agents use it to get permission to use a service.
- Checked by: OAuth authorization-server metadata
- OAuth authorization-server metadata Agent readiness
- A file (RFC 8414) describing a login server's endpoints and capabilities, so an agent can learn how to authenticate without being hard-coded.
- Checked by: OAuth authorization-server metadata
- OAuth protected-resource metadata Agent readiness
- A file (RFC 9728) in which a protected API or MCP server says which authorization server guards it. It is the first step of the MCP sign-in handshake.
- Checked by: OAuth protected-resource metadata
- Open Graph The page
- A set of meta tags (og:title, og:description, og:image, og:url) that control the preview card when a link is shared in social apps, chat apps and some AI interfaces.
- Checked by: Open Graph tags for sharing
- OpenID Connect discovery Agent readiness
- The standard file at /.well-known/openid-configuration that tells clients where a sign-in provider's token, authorisation and key endpoints are.
- Checked by: OpenID Connect discovery
- Organization (schema) Structured data
- The schema.org type for a business or organisation: its name, logo, and links to its other profiles. It is how a machine learns who stands behind a site.
- Checked by: Organization schema with sameAs
P
- Page weight Performance
- The total size of everything a page downloads: HTML, images, scripts, styles and fonts.
- Checked by: Images are optimised
- PageSpeed Insights (PSI) Performance
- Google's free service that reports a page's field data and runs a Lighthouse lab test. Our Performance card is powered by it.
- Checked by: Largest Contentful Paint (LCP)
- Partial parse / malformed markup Our own words
- When a page is too large or broken to parse completely, the checks that depend on its full body are marked "couldn't judge", and a notice says which kind were skipped.
- Pass Our own words
- We looked, and the page does what good looks like.
- Performance card Our own words
- The separate card showing PageSpeed's lab score and Core Web Vitals for the page. Never part of the headline. Replays show what was measured at the time; the on-demand control runs a fresh single-page scan.
- Checked by: Page weight and render-blocking
- Permalink (/r/<id>) Our own words
- The permanent address of one scan, showing it exactly as it was, including the performance measured at the time. It never quietly re-scans.
- Phase (checklist) Checklist vocabulary
- One of the checklist's stages, in order from strategy to getting cited by AI answer engines.
- Read more about Phase (checklist)
- Pinch-to-zoom (WCAG 1.4.4) The page
- Letting people enlarge a page on a phone. Blocking it fails an accessibility success criterion.
- Checked by: Mobile viewport is set
- Pinned pages Our own words
- Pages you choose to include in every future scan of a site, alongside the sample we pick.
- Project / Your sites Our own words
- Your list of site-scanned sites, and one site's history of runs, score trend and what changed. Needs an account.
- Public page history ("This page over time") Our own words
- Every scan of the same address, by anyone, kept as a dated history. It belongs to the page, not to the person who ran it.
Q
- QAPage Structured data
- The schema.org type for a page built around one question that users answer, like a forum thread. It is a different shape from an FAQ page.
- Checked by: FAQ / Q&A schema
R
- Rate limit Bots and access
- A cap on how many requests are accepted in a short period. Sites use one against heavy crawlers (HTTP 429); we use one on the scanner.
- Checked by: Automated visitors can reach your pages
- Redirect / 301 vs 302 / 307 The page
- A server answer that sends the visitor to another address. A 301 (or 308) says the move is permanent; a 302 or 307 says it is temporary.
- Checked by: Reached without a redirect chain
- Redirect chain The page
- A redirect that leads to another redirect before reaching the page. Each hop is an extra request, and some clients give up.
- Checked by: Reached without a redirect chain
- Reduced motion (prefers-reduced-motion) Checklist vocabulary
- A system setting people use to ask for less animation. A site can detect it and tone movement down.
- Referrer-Policy Security
- A header that controls how much of the current address is passed on to the next site when a visitor follows a link.
- Checked by: Core security headers are set
- rel="describedby" Files and signals
- A link relation meaning "this other document describes this one". It is how a site can point machines at its llms.txt.
- Checked by: Resource-discovery Link headers
- Render-blocking resource Performance
- A script or stylesheet the browser must download before it can show anything.
- Checked by: Page weight and render-blocking
- Responsive images (srcset) Performance
- Offering several sizes of one image so each device downloads the one it needs.
- Checked by: Images are optimised
- Rich result The field
- A search result with extra visual features (stars, breadcrumbs, FAQs) that a search engine builds from structured data. It is a search-display feature, separate from whether an AI can read the markup.
- Checked by: Structured data (JSON-LD) is present
- robots.txt Bots and access
- A plain-text file at the root of a site that tells crawlers, by user agent name, which parts they may fetch. It is a published standard (RFC 9309), and it is a request, not a lock.
- Checked by: robots.txt is present and not over-blocking
S
- sameAs Structured data
- A schema.org property listing a thing's other official profiles (LinkedIn, Wikipedia, social accounts), so engines can confirm those profiles are the same entity.
- Checked by: Organization schema with sameAs
- Scan / single-page scan Our own words
- One check of one address: we fetch the page the way a machine would, run every check, and report. Free, anonymous, about 20 seconds.
- Schema-content parity Structured data
- The principle that every fact in structured data should also be visible on the page. Markup that says something the page does not can be distrusted.
- Checked by: Structured data is valid
- schema.org Structured data
- The shared vocabulary of types (Organization, Article, Product, FAQPage…) and properties that search engines and AI tools recognise in structured data. It was founded by Google, Microsoft, Yahoo and Yandex.
- Checked by: Author / publisher is clear Read more about schema.org
- Score / headline score (0 to 100) Our own words
- One number out of 100 for how machine-readable the page is, built from the scored categories only. Performance, Advanced and Good practice never enter it.
- Scoring version Our own words
- A stamp on every scan and site run naming which version of our scoring rules produced it, so history never shows a change our rules caused as a change the site made.
- Search engine / SERP (search engine results page) The field
- A search engine is a service like Google or Bing that indexes the web and lists results. The SERP is the page of results it shows. We use Google as a contrast or a source on crawling, never as the yardstick for AI readability (Rule 6).
- Search engine crawler (Googlebot, Bingbot) Bots and access
- The crawlers that build Google's and Bing's search indexes. They render JavaScript, which most AI crawlers do not.
- Checked by: robots.txt is present and not over-blocking
- Search intent Checklist vocabulary
- What a person searching a query actually wants: to learn, to compare, to buy, or to reach a specific site.
- Read more about Search intent
- Security headers Security
- HTTP headers that tell browsers to apply protections: HSTS, X-Content-Type-Options, a framing policy and a Referrer-Policy.
- Checked by: Core security headers are set
- security.txt Files and signals
- A text file at /.well-known/security.txt (RFC 9116) telling security researchers how to report a vulnerability. It needs a contact and an expiry date to be trusted.
- Checked by: security.txt is published
- Self-referencing canonical The page
- A canonical that points at the page's own address, which is what a normal page should declare.
- Checked by: This page declares its canonical URL
- Semantic HTML / semantic landmarks The page
- HTML elements that say what a region is (header, nav, main, article, footer) instead of generic div boxes. Machines and screen readers use them to find the main content and skip the chrome.
- Checked by: Semantic HTML landmarks
- Sent back to the kitchen (verdict, bottom tier) Our own words
- "AI engines can barely get a bite of this one — time for a full recipe rework."
- SEO (search-engine optimisation) The field
- Making pages easy for traditional search engines to crawl, index and show in their results. Bot Appétit checks the SEO fundamentals that machines still depend on, not rankings.
- Server version leakage (Server, X-Powered-By) Security
- Response headers that announce the exact software and version a site runs, which gives attackers a head start.
- Checked by: No server version leakage
- Server-side rendering (server-rendered) The page
- Building the page's HTML on the server, so the content is already in the response before any script runs. Machines that don't run JavaScript can read it.
- Checked by: Your answer is in the HTML
- Signatures directory / JWKS (JSON Web Key Set) Bots and access
- The public file where a bot operator lists its signing keys, at /.well-known/http-message-signatures-directory, in the JWKS format. Only a domain that operates a bot needs one.
- Checked by: Web Bot Auth (operates a signed bot)
- Signed bot / signed request Bots and access
- A bot whose requests carry an HTTP message signature a site can verify. Bot Appétit's scanner is one.
- Checked by: Web Bot Auth (operates a signed bot)
- Site scan Our own words
- A scan of a representative sample of a whole site, each sampled page getting its own full page scan, rolled up into a site report. Free with an account.
- Site structure / click depth Checklist vocabulary
- How pages are grouped and linked, and how many clicks a page is from the homepage.
- Read more about Site structure / click depth
- Site-level check / host fact Our own words
- A check about the whole host rather than one page (robots.txt, sitemap, llms.txt, security.txt, www vs non-www and similar). In a site scan it is answered once for the host.
- Checked by: robots.txt is present and not over-blocking
- Sitemap index Files and signals
- A sitemap that lists other sitemaps, used by large sites. The site scan follows them.
- Checked by: XML sitemap is present and referenced
- Snippet The page
- The short description a search engine shows under a result's title, often taken from the meta description.
- Checked by: Compelling meta description
- Speed Index Performance
- A lab measure of how quickly the visible part of the page fills in.
- Checked by: Page weight and render-blocking
- Staging subdomain (*.webflow.io) Checklist vocabulary
- The free address Webflow gives a site before and alongside its real domain. If it can be indexed, it duplicates the real site.
- Checked by: Not accidentally set to noindex
- Structured data Structured data
- Facts about a page written in a fixed vocabulary a machine can read directly (what the page is, who published it, what it sells), separate from the prose.
T
- Tap target Checklist vocabulary
- A button or link on a touch screen. It needs to be large enough, and far enough from its neighbours, for a finger.
- TBT (Total Blocking Time) Performance
- In a lab test, how long the page's main thread was too busy to respond to input. It stands in for INP, which a lab test cannot measure.
- Checked by: Interaction to Next Paint (INP)
- Template sampling ("The shape of your site") Our own words
- Grouping a site's URLs by page template and scanning a sample of each, so a finding can say how many pages it affects.
- The checklist (Webflow SEO & AEO Checklist) Our own words
- Our free pre-launch SEO and AEO checklist for Webflow sites. It is related reading, not a fix for a failed check.
- Read more about The checklist (Webflow SEO & AEO Checklist)
- Title tag (<title>) The page
- The page's name in the HTML head. It is the first thing a search engine or answer engine reads to know what the page is, and it shows in browser tabs and results.
- Checked by: Has a title tag
- To fix Our own words
- We looked, and the thing is missing or broken.
- Trailing slash The page
- Whether an address ends in /. A site should settle on one form so each page has one address.
- Checked by: This page declares its canonical URL
- TTFB (Time to First Byte) Performance
- How long the server takes to start sending the page after it is asked. A supporting metric.
- Checked by: Page weight and render-blocking
- Twitter / X card The page
- X's own preview-card tags. When they are missing, X falls back to Open Graph.
- Checked by: Twitter/X card tags
U
- UCP (Universal Commerce Protocol) Agent readiness
- An open standard co-developed by Google for AI shopping agents: a JSON manifest at /.well-known/ucp declares what a business sells and how to transact. We only check it on sites that sell something.
- Checked by: Universal Commerce Protocol manifest
- Undercooked (verdict, fourth tier) Our own words
- "There's a lot on the menu bots can't taste yet — worth a proper pass."
- Checked by: AI crawlers aren't blocked
- Unused CSS Checklist vocabulary
- Style rules shipped to the browser that no element uses, adding weight for nothing.
- Checked by: Page weight and render-blocking
- URL slug The page
- The readable last part of a page's address, such as /pricing/.
- Checked by: Topic is consistent across title, H1 and description
- User agent Bots and access
- The name a browser or bot sends with every request to say what it is, for example GPTBot or our own BotAppetitBot/1.0. robots.txt rules are written against these names.
- Checked by: AI crawlers aren't blocked
V
- Validation (of structured data) Structured data
- Checking that each block parses as JSON and declares what it describes. Broken blocks are ignored by machines.
- Checked by: Structured data is valid
- Verdict Our own words
- The kitchen-themed label that goes with the score, one of five tiers below.
- Verdict ceiling (AI-blocked cap) Our own words
- When a site's robots.txt turns the major AI crawlers away, the verdict is held down whatever the score, and a notice on the report says why: a site AI can't read isn't ready for AI, however tidy its markup.
- Checked by: AI crawlers aren't blocked
- Verified bot Bots and access
- A crawler a bot-protection vendor has confirmed really belongs to the company it names, so it can be let through where unknown bots are not.
- Checked by: Automated visitors can reach your pages
- Viewport meta tag The page
- A tag that tells phones to lay the page out at the screen's real width instead of a pretend desktop width. Without it, responsive layouts never switch on.
- Checked by: Mobile viewport is set
- Vulnerability disclosure Security
- Reporting a security flaw privately to the site owner so it can be fixed. security.txt says where to send those reports.
- Checked by: security.txt is published
W
- WCAG Checklist vocabulary
- The Web Content Accessibility Guidelines, the W3C standard for accessible web content.
- Checked by: Mobile viewport is set
- Weak-course advisory Our own words
- A callout when the headline looks healthy but one scored category is not.
- Web Bot Auth Bots and access
- An IETF Internet-Draft for bots to prove who they are: the bot operator publishes signing keys on its own domain, signs each request, and a site can verify the signature. It replaces guessing from IP addresses and user agent names.
- Checked by: Web Bot Auth (operates a signed bot) Read more about Web Bot Auth
- Web fonts / font loading Checklist vocabulary
- Custom typefaces a page downloads. Each family and weight adds to the page's size and can delay text appearing.
- Checked by: Page weight and render-blocking
- WebMCP Agent readiness
- A proposed browser standard that lets a web page register tools an AI agent running in the browser can call directly, instead of scraping the page. It is a Draft Community Group Report from the W3C Web Machine Learning Community Group, not a W3C Standard. We do not check for it.
- WebSite / WebPage (structural types) Structured data
- Generic schema.org types that say "this is a site" or "this is a page" and little else.
- Checked by: Right schema type for the page
- /.well-known/ Files and signals
- A reserved folder at the root of a site where standards put machine-readable files at predictable addresses, so a program knows where to look without asking. security.txt, the Web Bot Auth directory and most agent-readiness files live there.
- Checked by: security.txt is published
- Wildcard rule (User-agent: *) Bots and access
- The robots.txt group that applies to every crawler not named elsewhere in the file. A site that only has the wildcard has made no bot-by-bot decision.
- Checked by: AI bot rules are explicit and intentional
- Worth a look Our own words
- We looked and found something partial, optional or worth a second look. Not "a thing to fix".
X
- X-Content-Type-Options Security
- A header telling browsers not to guess a file's type, which blocks a class of attacks.
- Checked by: Core security headers are set
- X-Frame-Options / CSP frame-ancestors Security
- Two ways of saying which other sites may show this page inside a frame, which defends against clickjacking.
- Checked by: Core security headers are set
- X-Robots-Tag The page
- The same robots directives as the meta tag, sent as an HTTP header instead. It also works for files that aren't HTML, such as PDFs and images.
- Checked by: Not accidentally set to noindex
- XML sitemap Files and signals
- A file listing a site's URLs in a standard XML format, so crawlers can discover pages without following every link. It is usually at /sitemap.xml and referenced from robots.txt.
- Checked by: XML sitemap is present and referenced
Y
- Your structured data (viewer) Our own words
- The report panel showing the JSON-LD we found on your page, with a copy button. It shows your own markup; it never writes new markup for you.
- Checked by: Structured data (JSON-LD) is present
See the words on your own page
A free scan runs all 66 checks in about 20 seconds — no signup. Everything we check