Ask an AI assistant the same question twice and you will often get two different source lists. In January 2026 SparkToro and Gumshoe had volunteers run 12 prompts about 3,000 times across ChatGPT, Claude and Google's AI. The same list of recommendations came back less than once in 100 runs, and in the same order closer to once in 1,000. Yet some brands appeared in 60–90% of answers for a given intent (Search Engine Land, 28 January 2026). Individual answers are noisy; how often you appear across many answers is stable enough to measure. So tracking AI citations means sampling: you don't record whether you were cited, you estimate how often.
Disclosure: citelity, which I build, tracks citations. Where it fits and where other tools fit better is spelled out below.
What the engines report directly
Less than you'd hope, more than a year ago.
- Bing Webmaster Tools, AI Performance (public preview since February 2026): total citations, cited pages, the "grounding queries" used to find your content, and citations per URL, across Copilot, Bing's AI summaries and selected partner integrations. No clicks or engagement data (Koozai). It is the only first-party report that counts citations.
- Google Search Console, generative AI report (all sites since 31 August 2026): impressions in AI Overviews, AI Mode and Discover AI features by page, country, device and date. No clicks, CTR, position or queries (Search Engine Journal).
- Google Analytics 4 added an "AI Assistant" default channel on 14 May 2026, assigning the medium
ai-assistantto referrals from recognized assistants including ChatGPT, Gemini and Claude (Search Engine Roundtable). That counts clicks, not citations. - ChatGPT, Perplexity, Claude and the Gemini app offer publishers nothing. For those, you sample.
If your GA4 property predates the new channel or you want your own grouping, a custom channel rule on session source does the same job:
^(chatgpt\.com|chat\.openai\.com|(www\.)?perplexity\.ai|gemini\.google\.com|claude\.ai|copilot\.microsoft\.com)$
How many runs make a citation rate trustworthy?
More than most people use. If you run a prompt ten times and get cited seven times, the honest statement is not "70%". It is "somewhere between about 40% and 89%". The table shows the 95% range for the true rate at different sample sizes:
| Runs | Times cited | Plausible true rate (95% range) |
|---|---|---|
| 2 | 1 | 9% – 91% |
| 5 | 4 | 38% – 96% |
| 10 | 7 | 40% – 89% |
| 10 | 3 | 11% – 60% |
| 20 | 14 | 48% – 85% |
| 50 | 35 | 56% – 81% |
| 100 | 50 | 40% – 60% |
| 100 | 30 | 22% – 40% |
Wilson score intervals, standard binomial arithmetic. They assume runs are independent; runs from the same day or session are correlated, so the real ranges are wider.
Three things follow.
A single prompt's rate is only good for big differences. Ten runs can tell "usually cited" from "almost never cited". They cannot tell 50% from 70%.
Detecting the effect of a change needs dozens of observations on each side. To have an 80% chance of spotting a real lift at the usual 5% significance level, you need roughly this many observations before and again after:
| True change | Observations per period |
|---|---|
| 30% → 60% | ~42 |
| 10% → 30% | ~62 |
| 30% → 50% | ~93 |
| 50% → 70% | ~93 |
| 40% → 50% | ~390 |
So pool prompts into a portfolio rate. Twenty prompts run five times each gives 100 observations per period, enough to detect a 20-point lift in your overall citation rate. It won't tell you which single prompt improved, and pooling assumes the prompts behave broadly alike, so keep each portfolio to one topic or page group. Per-prompt numbers stay useful for spotting zeros: a prompt with 0 out of 10 is almost certainly below 30%.
Spread the runs. Five runs in five minutes are closer to one observation than to five. Run across at least three days, and keep the session type (logged in or not, location) consistent between the before and after periods.
When to throw your baseline away
When the engine changes underneath you. A model swap moves citations far more than ordinary day-to-day noise:
- Gemini 3 in AI Overviews, 27 January 2026. SE Ranking found 42.4% of previously cited domains replaced across 100,000 keywords.
- Google's AI Overview and AI Mode link redesign, early May 2026. Inline links next to the supporting text changed what "position" means (Search Engine Land).
- GPT-5.5 in ChatGPT, 22–23 May 2026. SISTRIX saw 47% of citation patterns change overnight in its German dataset, against a normal day-to-day variation of 1–2% (Search Engine Journal).
A before/after comparison that straddles one of these measures the model change, not your edit. Note the dates of model releases in your tracking sheet and start a new baseline after each one.
Ways to collect the samples
By hand. A spreadsheet with one row per run: date, engine, prompt, session type, cited (yes/no), position, the quoted sentence. Fifteen prompts on four engines at ten runs each is 600 rows, roughly five to ten hours at 30–60 seconds a run. Fine for a baseline, too slow for weekly monitoring.
Through official APIs. Four engines return cited URLs programmatically:
- Perplexity's Sonar API returns citations with each answer.
- OpenAI's web search tool returns
url_citationannotations and, optionally, the full list of sources consulted. - Gemini's Grounding with Google Search returns cited URLs and the search queries it ran.
- Claude's API has a web search tool that returns cited sources.
The caveat applies to all of them: API answers can differ from what people see in the consumer apps (different model settings, no personalization, different query rewriting). Keep API data labeled as API data.
Through SERP data providers. AI Overviews have no official API. Paid SERP APIs such as DataForSEO and SerpApi return the Overview and its sources and are far more stable than the open-source scrapers on GitHub, which break whenever Google changes its markup.
What the paid tools cost
Prices from each vendor's own page, checked 30 September 2026:
| Tool | Entry plan | Higher tiers | Notes |
|---|---|---|---|
| Otterly.AI | Lite $29/mo, 15 prompts | Standard $189 (100 prompts), Premium $489 (400) | Annual billing −15%. Claude, Gemini and AI Mode are paid add-ons |
| AIclicks | Starter $59/mo, 30 prompts, 3 models | Pro $189 (150 prompts, 4 models), Business $499 (300, 6 models) | 3-day trial |
| LLMrefs | $79/mo, single plan | n/a | 500 prompts, 8 engines, 7-day trial |
| Peec AI | Starter €85/mo, 50 prompts | Pro €205 (150), Advanced €425 (350) | Priced in euros; 3 models of your choice |
| Profound | 7-day free trial | Enterprise, custom pricing | Self-serve agency plan listed without a price |
Price per prompt is the easy comparison: $1.89 a prompt on Otterly Standard, $1.26 on AIclicks Pro, about $0.16 on LLMrefs. The comparison that matters more is runs per prompt. Given the table above, a tool that checks each prompt once a week needs months before a single prompt's rate means much, though its portfolio rate can be useful within weeks. Ask every vendor:
- How many times is each prompt run per reporting period, and when?
- Which engines come from official APIs, which from live app sampling, and which from SERP data? Are they labeled?
- Is the quoted sentence captured, or only the URL?
- Are model-release dates marked on the trend charts?
citelity (mine). $49, $99 or $199 a month, tracking 25, 30 or 50 prompts per scan across ChatGPT, Google AI Overviews, Gemini, Claude and Perplexity. Each weekly scan runs every tracked prompt once per engine, so per-prompt rates build up over weeks while the portfolio rate is usable sooner. Gemini, Claude and Perplexity are sampled through their official APIs, ChatGPT through live samples and AI Overviews through search data, and each result carries a badge saying which. Tracking sits inside a loop that starts from Search Console pages losing clicks, writes the fix and measures the result. If you only need tracking and already have a content process, a monitoring-only tool will be cheaper.
A tracking setup for one site
- Pick one portfolio. 15–20 prompts about one topic or page group, phrased the way people ask. Include questions that need current or specific information; broad definitions often get no sources at all.
- Choose two engines. AI Overviews plus the one your audience uses most. Perplexity gives you the most data per run because it cites on nearly every answer.
- Baseline. Five runs per prompt per engine over at least three days: about 100 observations per engine.
- Change one thing. Rewrite the pages behind the portfolio, then wait for them to be recrawled (check your server logs for the engines' bots).
- Re-measure the same way and compare portfolio rates using the second table above. If a model release fell in between, re-baseline instead of comparing.
- Write the method down (prompts, engines, run counts, session type) so the next measurement is comparable.
The diagnostic for a single citation that disappeared is in that guide; engine-specific access checks are in the guides for ChatGPT, Perplexity, AI Overviews, Gemini and Claude.
A tool says I'm cited but I can't reproduce it manually. Who's right?
Should I track position in the source list?
Sources
- Search Engine Land — SparkToro/Gumshoe: AI recommendation lists rarely repeat, 28 Jan 2026
- Koozai — Bing Webmaster Tools launches AI Performance reporting, 13 Feb 2026
- Search Engine Journal — Search Console AI reports rolled out worldwide, 31 Aug 2026
- Search Engine Roundtable — Google Analytics adds AI Assistant traffic channel, 14 May 2026
- SE Ranking — Gemini 3's impact on AI Overviews, 26 Feb 2026
- Search Engine Land — Google updates links within AI Overviews and AI Mode, 6 May 2026
- Search Engine Journal — ChatGPT citations changed after GPT-5.5 (SISTRIX data)
- OpenAI — Web search tool documentation
- Google AI for Developers — Grounding with Google Search
- Vendor pricing pages: Otterly.AI, AIclicks, LLMrefs, Peec AI, Profound
Updated October 1, 2026: refocused on sampling and sample sizes; added Bing's AI Performance report, GA4's AI Assistant channel and the official API routes for ChatGPT, Gemini and Claude; corrected Peec and LLMrefs pricing; replaced the claim that no volatility data exists with the SparkToro/Gumshoe study.