Tracking AI citations: how many runs make a rate you can trust

AI answers cite different sources from one run to the next, so tracking means estimating a rate. How many runs and prompts make that rate trustworthy, what the engines report directly, and what the tools cost.

· Updated 10 min read

Ask an AI assistant the same question twice and you will often get two different source lists. In January 2026 SparkToro and Gumshoe had volunteers run 12 prompts about 3,000 times across ChatGPT, Claude and Google's AI. The same list of recommendations came back less than once in 100 runs, and in the same order closer to once in 1,000. Yet some brands appeared in 60–90% of answers for a given intent (Search Engine Land, 28 January 2026). Individual answers are noisy; how often you appear across many answers is stable enough to measure. So tracking AI citations means sampling: you don't record whether you were cited, you estimate how often.

Disclosure: citelity, which I build, tracks citations. Where it fits and where other tools fit better is spelled out below.

What the engines report directly

Less than you'd hope, more than a year ago.

  • Bing Webmaster Tools, AI Performance (public preview since February 2026): total citations, cited pages, the "grounding queries" used to find your content, and citations per URL, across Copilot, Bing's AI summaries and selected partner integrations. No clicks or engagement data (Koozai). It is the only first-party report that counts citations.
  • Google Search Console, generative AI report (all sites since 31 August 2026): impressions in AI Overviews, AI Mode and Discover AI features by page, country, device and date. No clicks, CTR, position or queries (Search Engine Journal).
  • Google Analytics 4 added an "AI Assistant" default channel on 14 May 2026, assigning the medium ai-assistant to referrals from recognized assistants including ChatGPT, Gemini and Claude (Search Engine Roundtable). That counts clicks, not citations.
  • ChatGPT, Perplexity, Claude and the Gemini app offer publishers nothing. For those, you sample.

If your GA4 property predates the new channel or you want your own grouping, a custom channel rule on session source does the same job:

^(chatgpt\.com|chat\.openai\.com|(www\.)?perplexity\.ai|gemini\.google\.com|claude\.ai|copilot\.microsoft\.com)$

How many runs make a citation rate trustworthy?

More than most people use. If you run a prompt ten times and get cited seven times, the honest statement is not "70%". It is "somewhere between about 40% and 89%". The table shows the 95% range for the true rate at different sample sizes:

RunsTimes citedPlausible true rate (95% range)
219% – 91%
5438% – 96%
10740% – 89%
10311% – 60%
201448% – 85%
503556% – 81%
1005040% – 60%
1003022% – 40%

Wilson score intervals, standard binomial arithmetic. They assume runs are independent; runs from the same day or session are correlated, so the real ranges are wider.

Three things follow.

A single prompt's rate is only good for big differences. Ten runs can tell "usually cited" from "almost never cited". They cannot tell 50% from 70%.

Detecting the effect of a change needs dozens of observations on each side. To have an 80% chance of spotting a real lift at the usual 5% significance level, you need roughly this many observations before and again after:

True changeObservations per period
30% → 60%~42
10% → 30%~62
30% → 50%~93
50% → 70%~93
40% → 50%~390

So pool prompts into a portfolio rate. Twenty prompts run five times each gives 100 observations per period, enough to detect a 20-point lift in your overall citation rate. It won't tell you which single prompt improved, and pooling assumes the prompts behave broadly alike, so keep each portfolio to one topic or page group. Per-prompt numbers stay useful for spotting zeros: a prompt with 0 out of 10 is almost certainly below 30%.

Spread the runs. Five runs in five minutes are closer to one observation than to five. Run across at least three days, and keep the session type (logged in or not, location) consistent between the before and after periods.

When to throw your baseline away

When the engine changes underneath you. A model swap moves citations far more than ordinary day-to-day noise:

  • Gemini 3 in AI Overviews, 27 January 2026. SE Ranking found 42.4% of previously cited domains replaced across 100,000 keywords.
  • Google's AI Overview and AI Mode link redesign, early May 2026. Inline links next to the supporting text changed what "position" means (Search Engine Land).
  • GPT-5.5 in ChatGPT, 22–23 May 2026. SISTRIX saw 47% of citation patterns change overnight in its German dataset, against a normal day-to-day variation of 1–2% (Search Engine Journal).

A before/after comparison that straddles one of these measures the model change, not your edit. Note the dates of model releases in your tracking sheet and start a new baseline after each one.

Ways to collect the samples

By hand. A spreadsheet with one row per run: date, engine, prompt, session type, cited (yes/no), position, the quoted sentence. Fifteen prompts on four engines at ten runs each is 600 rows, roughly five to ten hours at 30–60 seconds a run. Fine for a baseline, too slow for weekly monitoring.

Through official APIs. Four engines return cited URLs programmatically:

  • Perplexity's Sonar API returns citations with each answer.
  • OpenAI's web search tool returns url_citation annotations and, optionally, the full list of sources consulted.
  • Gemini's Grounding with Google Search returns cited URLs and the search queries it ran.
  • Claude's API has a web search tool that returns cited sources.

The caveat applies to all of them: API answers can differ from what people see in the consumer apps (different model settings, no personalization, different query rewriting). Keep API data labeled as API data.

Through SERP data providers. AI Overviews have no official API. Paid SERP APIs such as DataForSEO and SerpApi return the Overview and its sources and are far more stable than the open-source scrapers on GitHub, which break whenever Google changes its markup.

AI
Free tool · No signup
Free AI Prompt Demand Checker
Check real monthly search demand for the query behind a prompt — with the data source named, not estimated.
Check prompt demand →

What the paid tools cost

Prices from each vendor's own page, checked 30 September 2026:

ToolEntry planHigher tiersNotes
Otterly.AILite $29/mo, 15 promptsStandard $189 (100 prompts), Premium $489 (400)Annual billing −15%. Claude, Gemini and AI Mode are paid add-ons
AIclicksStarter $59/mo, 30 prompts, 3 modelsPro $189 (150 prompts, 4 models), Business $499 (300, 6 models)3-day trial
LLMrefs$79/mo, single plann/a500 prompts, 8 engines, 7-day trial
Peec AIStarter €85/mo, 50 promptsPro €205 (150), Advanced €425 (350)Priced in euros; 3 models of your choice
Profound7-day free trialEnterprise, custom pricingSelf-serve agency plan listed without a price

Price per prompt is the easy comparison: $1.89 a prompt on Otterly Standard, $1.26 on AIclicks Pro, about $0.16 on LLMrefs. The comparison that matters more is runs per prompt. Given the table above, a tool that checks each prompt once a week needs months before a single prompt's rate means much, though its portfolio rate can be useful within weeks. Ask every vendor:

  1. How many times is each prompt run per reporting period, and when?
  2. Which engines come from official APIs, which from live app sampling, and which from SERP data? Are they labeled?
  3. Is the quoted sentence captured, or only the URL?
  4. Are model-release dates marked on the trend charts?

citelity (mine). $49, $99 or $199 a month, tracking 25, 30 or 50 prompts per scan across ChatGPT, Google AI Overviews, Gemini, Claude and Perplexity. Each weekly scan runs every tracked prompt once per engine, so per-prompt rates build up over weeks while the portfolio rate is usable sooner. Gemini, Claude and Perplexity are sampled through their official APIs, ChatGPT through live samples and AI Overviews through search data, and each result carries a badge saying which. Tracking sits inside a loop that starts from Search Console pages losing clicks, writes the fix and measures the result. If you only need tracking and already have a content process, a monitoring-only tool will be cheaper.

A tracking setup for one site

  1. Pick one portfolio. 15–20 prompts about one topic or page group, phrased the way people ask. Include questions that need current or specific information; broad definitions often get no sources at all.
  2. Choose two engines. AI Overviews plus the one your audience uses most. Perplexity gives you the most data per run because it cites on nearly every answer.
  3. Baseline. Five runs per prompt per engine over at least three days: about 100 observations per engine.
  4. Change one thing. Rewrite the pages behind the portfolio, then wait for them to be recrawled (check your server logs for the engines' bots).
  5. Re-measure the same way and compare portfolio rates using the second table above. If a model release fell in between, re-baseline instead of comparing.
  6. Write the method down (prompts, engines, run counts, session type) so the next measurement is comparable.

The diagnostic for a single citation that disappeared is in that guide; engine-specific access checks are in the guides for ChatGPT, Perplexity, AI Overviews, Gemini and Claude.

A tool says I'm cited but I can't reproduce it manually. Who's right?
Possibly both. With a handful of runs on each side, a tool reporting 6 of 10 and you seeing 2 of 10 are not clearly different rates. Run the prompt 20 times yourself across several days and compare ranges, not single numbers. Also ask whether the tool samples the consumer app or an API, and from which location, because those give different answers.
Should I track position in the source list?
Only loosely. The SparkToro and Gumshoe test found the same list in the same order roughly once in 1,000 runs, and Google's inline-link redesign means AI Overview sources no longer have one clear order. Presence across runs is the measurement that holds up; record position, but don't read meaning into moves of one or two places.

Sources

Updated October 1, 2026: refocused on sampling and sample sizes; added Bing's AI Performance report, GA4's AI Assistant channel and the official API routes for ChatGPT, Gemini and Claude; corrected Peec and LLMrefs pricing; replaced the claim that no volatility data exists with the SparkToro/Gumshoe study.