AEO content checklist: 47 checks, from crawler access to quarterly upkeep

47 AEO checks in the order you do them: retrievability (now including Search Console's AI control), planning, structure, schema, authorship, validation, upkeep.

· Updated 10 min read

Forty-seven checks, in the order you'd actually do them: 6 for retrievability, 8 for planning, 10 for on-page structure, 7 for schema, 6 for authorship, 5 for validation and 5 for quarterly upkeep. Use the whole list on new articles; for existing pages, start with your top 20 by traffic.

Most checklists start at the writing stage. This one starts earlier, because a page that no retrieval crawler can fetch (or that Search Console has been told to keep out of AI features) is invisible no matter how well it's written, and that's usually an accident nobody notices.

Stage 1: retrievability (6 checks)

Do these once per site, then again after any deploy, CDN change or Search Console settings change.

  • OAI-SearchBot is allowed in robots.txt. This is what ChatGPT search uses; GPTBot is for training. ChatGPT-User fetches pages when a user asks about them, and OpenAI says robots.txt may not apply to it
  • PerplexityBot is allowed. Many 2024-era "block AI" templates block it along with the training bots. (Perplexity-User generally ignores robots.txt anyway)
  • Claude-SearchBot and Claude-User are allowed. ClaudeBot is training only; Anthropic says blocking Claude-User may reduce your visibility in user-directed search
  • Googlebot and Bingbot are allowed. Googlebot serves Search, AI Overviews and AI Mode together, with no separate robots.txt token for the AI features; Microsoft Copilot draws on Bing
  • Search Console's Search generative AI control is not set to opt out (Settings, then Search generative AI). Since August 31, 2026 it's available to every site, and switching it on removes the site from AI Overviews, AI Mode and Discover's AI features even with Googlebot allowed
  • CDN and WAF rules checked, and server logs show visits from each retrieval crawler in the last 30 days. Cloudflare has blocked AI crawlers by default on new domains since July 2025, independently of your robots.txt

The free AI Crawler Checker reads your robots.txt for the first four checks. It can't see your CDN rules or your Search Console setting, so those two stay manual.

AI
Free tool · No signup
Free AI Crawler Checker
Paste a domain → see which AI crawlers your robots.txt allows and which it blocks, with training and retrieval bots told apart.
Check your crawlers →

Stage 2: planning (8 checks)

Before writing.

  • The page targets a question people actually ask, not a keyword bent into a question
  • You've run the query in an engine and confirmed it produces an answer with citations
  • The query is winnable: the sources currently cited are sites your size, not only Wikipedia, YouTube and major outlets
  • The article has an original angle: your data, your test, your analysis
  • Two or three named sources are lined up for the main claims
  • At least one number you can attribute and link to
  • The direct answer is drafted first: 50 to 80 words that stand on their own
  • The page fits an existing topic cluster rather than starting a new one

On length: there's no target. Ahrefs found both short and long content cited in AI Overviews, and no published study shows a word count that improves citation. Length should follow from covering the question.

Stage 3: on-page structure (10 checks)

While writing and editing. Most of the citation outcome is decided here.

Opening:

  • The first sentence answers the article's main question
  • The first 50 to 80 words are a complete answer by themselves
  • No preamble: nothing like "In this guide we'll explore"

Sections:

  • Each H2 reads as a question a reader would ask
  • Each section opens with its own answer, not with background
  • Content comes in self-contained blocks of roughly 40 to 120 words that make sense when lifted out
  • Comparisons use a table, since rows extract cleanly

Claims:

  • Every factual claim names its source, never "studies show"
  • Every statistic links to a primary source
  • No keyword stuffing, the method that did worst in the original GEO benchmark

Why per-section openings matter: in Growth Memo's study of 18,012 ChatGPT citations (reported by Search Engine Land), 44.2% came from the first 30% of a page's text and 24.7% from the final 30%. Material deep in a page is quoted less often, not never, so each section should be able to answer on its own.

Stage 4: schema (7 checks)

More than ten rich result types have been retired since 2023, two of which used to be on this list (FAQ and HowTo).

  • Article or BlogPosting present as JSON-LD
  • Required fields present: headline, description, datePublished, dateModified, image
  • author is a full Person object, not a bare string
  • The Person has name, a url pointing to a real bio page, sameAs links to profiles that exist, and a relevant jobTitle
  • publisher is an Organization object, set site-wide
  • BreadcrumbList present, positions numbered from 1
  • Schema is rendered server-side: check view-source, not the browser inspector

Stage 5: authorship (6 checks)

  • A real named author, not "Editorial Team" or "Admin"
  • A bio page with substantive, relevant background
  • The bio lists specific experience: roles, years, prior work
  • The bio links to profiles that exist and describe the same person
  • A visible byline with a real photo near the top
  • At least one first-hand detail in the body, where it's true: what you tested, measured or saw

A named author with a presence elsewhere is also what connects mentions of a person or brand back to your pages, and mentions correlate more strongly with AI Overview visibility than backlinks do (the brand-signal data is in the AEO explainer).

Stage 6: validation (5 checks)

Right after publishing.

  • Rich Results Test: zero errors for the types that still produce rich results
  • Schema.org Validator: now the place to check FAQPage, since the Rich Results Test no longer does
  • Manual cross-check: the JSON-LD and the rendered page describe the same thing
  • Identity chain works: bio page loads, sameAs profiles exist, photo matches
  • Indexing requested through Search Console's URL Inspection

Stage 7: upkeep (5 checks, quarterly)

  • dateModified changed only when visible content changed
  • Statistics, prices and named tools re-verified
  • Schema still validates (theme and CMS updates break it quietly)
  • Internal links still resolve
  • Citation rate sampled on the target prompts and written down

This stage matters more than it looks. Seer Interactive's 2026 study of 47,097 citations across ChatGPT, Gemini and Perplexity found 75% of cited pages had been updated within a year and 88% within two. More to the point for a checklist: 72% of cited pages were fresh by update date but only 42% by publish date. Most of the freshness that gets cited comes from updating existing pages, not publishing new ones.

AI
Free tool · No signup
Free AEO Readiness Score
Paste content or a URL → a 0-100 score across ten AEO signals, with the fixes worth doing first.
Score your content →

If you only have time for ten

  1. Retrieval crawlers allowed in robots.txt
  2. CDN rules and the Search Console AI control checked
  3. A direct answer in the first 50 to 80 words
  4. An answer-first opening for every H2
  5. Self-contained blocks of roughly 40 to 120 words
  6. Every statistic linked to a named source
  7. Article schema with a full Person author
  8. A real bio page that the Person schema points to
  9. A table wherever the page compares things
  10. A recorded citation baseline before you change anything

The first two take about fifteen minutes and can invalidate everything else. The tenth is the one people skip, and without it you can't tell whether any of the rest worked.

Working through existing pages

Forty-seven checks across 50 pages is 2,350 checks, so order matters.

Batch by stage, not by page. Do the authorship checks across all 20 pages in one sitting, then schema across all 20. You stay in one tool and one frame of mind, which is usually quicker than finishing one page before starting the next.

Start with traffic at risk. The 20 pages that brought the most search traffic come first. In Search Console, the ones whose position held while clicks fell are the most urgent.

FAQ

How often should I re-run this on existing pages?
Run the upkeep stage quarterly on your top 20 pages by traffic, and the full list only when a page is substantially rewritten. Pages outside the top 20 rarely justify quarterly attention; once a year, or when traffic drops, is enough. Anything with prices, rankings or statistics in it needs checking more often, because stale facts are the first thing a reader or an engine will notice.
Does this checklist apply to product and category pages?
Stage 1 applies to every page, because a blocked crawler blocks everything. The rest is written for articles and guides. Product and category pages benefit from a short factual answer near the top, a comparison table and accurate Product markup, but question-shaped headings and long explanatory blocks usually don't fit them.

Next step

Do Stage 1 today, including the Search Console setting, because it takes fifteen minutes and protects every other hour. Then take three pages that already get traffic, run the whole list on them, and keep them as templates for the rest. The reasoning behind most of these checks is in what is AEO, schema detail in schema markup for AI Overviews, and measurement in how to track AI citations.

Sources

Updated October 1, 2026: added the Search Console AI control, Bingbot and the user-triggered fetchers to Stage 1 (merging the server-log check into the CDN check to keep 47), and replaced the freshness data with Seer's 2026 citation study.