Forty-seven checks, in the order you'd actually do them: 6 for retrievability, 8 for planning, 10 for on-page structure, 7 for schema, 6 for authorship, 5 for validation and 5 for quarterly upkeep. Use the whole list on new articles; for existing pages, start with your top 20 by traffic.
Most checklists start at the writing stage. This one starts earlier, because a page that no retrieval crawler can fetch (or that Search Console has been told to keep out of AI features) is invisible no matter how well it's written, and that's usually an accident nobody notices.
Stage 1: retrievability (6 checks)
Do these once per site, then again after any deploy, CDN change or Search Console settings change.
-
OAI-SearchBotis allowed in robots.txt. This is what ChatGPT search uses;GPTBotis for training.ChatGPT-Userfetches pages when a user asks about them, and OpenAI says robots.txt may not apply to it -
PerplexityBotis allowed. Many 2024-era "block AI" templates block it along with the training bots. (Perplexity-User generally ignores robots.txt anyway) -
Claude-SearchBotandClaude-Userare allowed.ClaudeBotis training only; Anthropic says blocking Claude-User may reduce your visibility in user-directed search -
GooglebotandBingbotare allowed. Googlebot serves Search, AI Overviews and AI Mode together, with no separate robots.txt token for the AI features; Microsoft Copilot draws on Bing - Search Console's Search generative AI control is not set to opt out (Settings, then Search generative AI). Since August 31, 2026 it's available to every site, and switching it on removes the site from AI Overviews, AI Mode and Discover's AI features even with Googlebot allowed
- CDN and WAF rules checked, and server logs show visits from each retrieval crawler in the last 30 days. Cloudflare has blocked AI crawlers by default on new domains since July 2025, independently of your robots.txt
The free AI Crawler Checker reads your robots.txt for the first four checks. It can't see your CDN rules or your Search Console setting, so those two stay manual.
Stage 2: planning (8 checks)
Before writing.
- The page targets a question people actually ask, not a keyword bent into a question
- You've run the query in an engine and confirmed it produces an answer with citations
- The query is winnable: the sources currently cited are sites your size, not only Wikipedia, YouTube and major outlets
- The article has an original angle: your data, your test, your analysis
- Two or three named sources are lined up for the main claims
- At least one number you can attribute and link to
- The direct answer is drafted first: 50 to 80 words that stand on their own
- The page fits an existing topic cluster rather than starting a new one
On length: there's no target. Ahrefs found both short and long content cited in AI Overviews, and no published study shows a word count that improves citation. Length should follow from covering the question.
Stage 3: on-page structure (10 checks)
While writing and editing. Most of the citation outcome is decided here.
Opening:
- The first sentence answers the article's main question
- The first 50 to 80 words are a complete answer by themselves
- No preamble: nothing like "In this guide we'll explore"
Sections:
- Each H2 reads as a question a reader would ask
- Each section opens with its own answer, not with background
- Content comes in self-contained blocks of roughly 40 to 120 words that make sense when lifted out
- Comparisons use a table, since rows extract cleanly
Claims:
- Every factual claim names its source, never "studies show"
- Every statistic links to a primary source
- No keyword stuffing, the method that did worst in the original GEO benchmark
Why per-section openings matter: in Growth Memo's study of 18,012 ChatGPT citations (reported by Search Engine Land), 44.2% came from the first 30% of a page's text and 24.7% from the final 30%. Material deep in a page is quoted less often, not never, so each section should be able to answer on its own.
Stage 4: schema (7 checks)
More than ten rich result types have been retired since 2023, two of which used to be on this list (FAQ and HowTo).
- Article or BlogPosting present as JSON-LD
- Required fields present:
headline,description,datePublished,dateModified,image -
authoris a full Person object, not a bare string - The Person has
name, aurlpointing to a real bio page,sameAslinks to profiles that exist, and a relevantjobTitle -
publisheris an Organization object, set site-wide - BreadcrumbList present, positions numbered from 1
- Schema is rendered server-side: check view-source, not the browser inspector
Stage 5: authorship (6 checks)
- A real named author, not "Editorial Team" or "Admin"
- A bio page with substantive, relevant background
- The bio lists specific experience: roles, years, prior work
- The bio links to profiles that exist and describe the same person
- A visible byline with a real photo near the top
- At least one first-hand detail in the body, where it's true: what you tested, measured or saw
A named author with a presence elsewhere is also what connects mentions of a person or brand back to your pages, and mentions correlate more strongly with AI Overview visibility than backlinks do (the brand-signal data is in the AEO explainer).
Stage 6: validation (5 checks)
Right after publishing.
- Rich Results Test: zero errors for the types that still produce rich results
- Schema.org Validator: now the place to check FAQPage, since the Rich Results Test no longer does
- Manual cross-check: the JSON-LD and the rendered page describe the same thing
- Identity chain works: bio page loads,
sameAsprofiles exist, photo matches - Indexing requested through Search Console's URL Inspection
Stage 7: upkeep (5 checks, quarterly)
-
dateModifiedchanged only when visible content changed - Statistics, prices and named tools re-verified
- Schema still validates (theme and CMS updates break it quietly)
- Internal links still resolve
- Citation rate sampled on the target prompts and written down
This stage matters more than it looks. Seer Interactive's 2026 study of 47,097 citations across ChatGPT, Gemini and Perplexity found 75% of cited pages had been updated within a year and 88% within two. More to the point for a checklist: 72% of cited pages were fresh by update date but only 42% by publish date. Most of the freshness that gets cited comes from updating existing pages, not publishing new ones.
If you only have time for ten
- Retrieval crawlers allowed in robots.txt
- CDN rules and the Search Console AI control checked
- A direct answer in the first 50 to 80 words
- An answer-first opening for every H2
- Self-contained blocks of roughly 40 to 120 words
- Every statistic linked to a named source
- Article schema with a full Person author
- A real bio page that the Person schema points to
- A table wherever the page compares things
- A recorded citation baseline before you change anything
The first two take about fifteen minutes and can invalidate everything else. The tenth is the one people skip, and without it you can't tell whether any of the rest worked.
Working through existing pages
Forty-seven checks across 50 pages is 2,350 checks, so order matters.
Batch by stage, not by page. Do the authorship checks across all 20 pages in one sitting, then schema across all 20. You stay in one tool and one frame of mind, which is usually quicker than finishing one page before starting the next.
Start with traffic at risk. The 20 pages that brought the most search traffic come first. In Search Console, the ones whose position held while clicks fell are the most urgent.
FAQ
How often should I re-run this on existing pages?
Does this checklist apply to product and category pages?
Next step
Do Stage 1 today, including the Search Console setting, because it takes fifteen minutes and protects every other hour. Then take three pages that already get traffic, run the whole list on them, and keep them as templates for the rest. The reasoning behind most of these checks is in what is AEO, schema detail in schema markup for AI Overviews, and measurement in how to track AI citations.
Sources
- OpenAI — Overview of OpenAI crawlers
- Perplexity — Perplexity crawlers
- Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Google — Common crawlers, including Google-Extended
- Search Engine Land — Search Console AI performance reports and Search generative AI control roll out globally
- MIT Technology Review — Cloudflare will now block AI bots by default, July 1, 2025
- Ahrefs — Short vs. long content in AI Overviews
- Search Engine Land — Growth Memo study of ChatGPT citations by page position, February 2026
- Google — FAQPage structured data documentation
- Google — AI features and your website
- Seer Interactive — Content recency's impact on AI visibility in 2026, July 2026
Updated October 1, 2026: added the Search Console AI control, Bingbot and the user-triggered fetchers to Stage 1 (merging the server-log check into the CDN check to keep 47), and replaced the freshness data with Seer's 2026 citation study.