citelity.Join waitlist →
June 29, 2026·14 min read·citation-tracking

How to track AI citations in 2026: methods that actually work

There's no dashboard for AI citations, and citation sets churn heavily. Here are the three real methods, what to measure, and which numbers are noise.

Tracking AI citations is manual work, and the unit of measurement is a rate rather than an event. No engine offers a Search Console equivalent, citation sets vary considerably between runs, and each engine exposes different amounts of data.

Three approaches exist: manual sampling (free, good to about 20 queries), semi-automated scripting (cheap, moderate setup), and paid tracking tools ($29–$500/month). This guide covers what each does, what to measure, and which numbers are signal.

Disclosure: I'm building citelity, which includes citation tracking. I've been specific below about what it does and where other tools fit better.

Why is AI citation tracking hard?

Three structural problems, and they compound.

No data is exposed. Google Search Console added a generative AI report in June 2026 showing AI Overview and AI Mode impressions — broken down by page, country, device and date. It contains no clicks, no CTR, no position, and no query-level citation data. For ChatGPT, Perplexity, Claude and standalone Gemini there is no publisher dashboard at all.

Citation sets churn. SE Ranking analysed 100,000 keywords before and after Google switched AI Overviews to Gemini 3 on 27 January 2026: 42.4% of previously cited domains were replaced and sources per answer rose 31.8%, from 11.55 to 15.22. No page changed — the model did. If a model swap can move that much, a single observation proves nothing.

Engines differ. Profound's dataset of over 680 million citations shows distinct source preferences per engine — Wikipedia leads ChatGPT citations at 7.8%, Reddit leads both Perplexity at 6.6% and Google AI Overviews at 2.2%. A tool that samples one engine well often handles the others poorly.

What should you measure?

Five metrics, in order of how much they tell you.

What's mostly noise: a single citation appearance; day-of-week patterns below 50 samples; position changes of one or two places; cross-engine comparisons on queries where one engine doesn't generate an AI answer at all.

Method 1: manual sampling

The right starting point for almost everyone. Free, scales to roughly 20 queries, and it teaches you what to look for before you spend anything.

Setup. A spreadsheet with columns for query, date, engine, session type (regular or incognito), cited yes/no, position if cited, and the extracted sentence if identifiable.

Routine. For each query: run it five to ten times across at least three different days, mixing regular and incognito sessions since personalisation can inflate appearance. Record each run as a row. After eight to ten runs, calculate the rate.

The arithmetic. Fifteen queries across four engines at ten runs each is 600 samples. At 30–60 seconds per run that's five to ten hours — feasible as a one-time baseline, not as weekly monitoring.

Where it stops working: past 20 queries, past two or three engines, or when you need to detect change within a week of an edit.

Method 2: semi-automated

Between fully manual and paid, there's a band that automates part of the work.

Perplexity API is the cleanest option. It returns citation data structurally, so a small script can run 20–50 queries weekly and log results to a database for a few dollars a month. It covers one engine well and nothing else.

Screenshot workflows — a full-page capture extension plus a rotating daily schedule — get you to roughly 25–35 queries weekly at about 30 minutes a day, with better records than pure manual.

Scraping scripts for AI Overviews exist on GitHub and break regularly as Google adjusts rendering. Budget for maintenance, or don't start.

ChatGPT Search and standalone Gemini have no clean API path for search-grounded responses. Treat both as manual-only surfaces.

A realistic stack: Perplexity API scripted weekly, manual sampling for AI Overviews and ChatGPT, spreadsheet aggregation. Roughly $5–20/month and two to four hours weekly.

Method 3: paid tracking tools

Current published pricing, which has moved considerably in this category over the past year.

ToolEntryNotes
Otterly.AI$29/moStandard $189, Premium $489. Cheapest genuine on-ramp
AIclicks$59/moPro $189, Business $499. Attribution-focused
LLMrefs$79/mo~500 prompts, up to 50 keywords, plus a free tier
Peec AI~$89/moTiered by prompt volume; strong competitive analysis
Profoundnot publishedDemo-led, enterprise custom. Largest public dataset

What to ask before buying:

Sampling depth. How many runs per prompt does the tool actually perform? A tool reporting citation status from one or two samples is reporting a coin flip. Five or more is the minimum for a meaningful rate. Few vendors publish this — ask.

Engine coverage. Does it cover the engines your audience uses? AI Overviews plus one other is the minimum useful pair for most content sites.

Snippet capture. Does it record which sentence was quoted? This is the most actionable detail and tools vary widely on whether they capture it.

Cost per tracked prompt. Divide monthly cost by prompt slots. A $189 plan covering 350 prompts is cheaper per unit than a $79 plan covering 50.

citelity (mine). Tracking sits inside a Find → Fix → Prove loop rather than standing alone: it starts from Search Console to find pages losing clicks while rankings hold, generates the fix, then measures the result in positions, citations and GA4 traffic. Tracked engines are ChatGPT, Perplexity, Claude and Google AI Overviews. $49 / $99 / $199 a month. If you only need tracking and already have a content process, a visibility-only tool is cheaper and I'd point you there.

How often should you sample?

AI
Free tool · No signup
Free AEO Readiness Score
Paste content or a URL → a 0-100 score across ten AEO signals, with the fixes worth doing first.
Score your content

Five mistakes that produce wrong conclusions

Claiming a win from one observation. You got cited this morning. That's one sample. Don't celebrate — or post about it — until you have a rate from five or more runs.

Comparing engines on incompatible queries. "Cited in Perplexity but not AI Overviews" might be a structural pattern, or the query might not trigger an AI Overview at all. Check that both engines actually generate an AI answer before drawing a conclusion.

Reading micro-position changes. Moving from cited source four to source three is within variance. Presence versus absence, and the rate, are the actionable measurements.

Sampling all at once. Ten queries in one hour gives you one session, not a pattern over time. Spread across at least three days.

Comparing against memory. "It feels like we were cited more last month" is not a baseline. If you didn't record it, you can't compare to it — and this is the single most common reason people conclude their AEO work isn't paying off.

A realistic 30-day setup

Week 1. Pick 10–15 queries you genuinely care about. Run ten samples each on AI Overviews plus one other engine. Record everything. Calculate baseline rates.

Week 2. Decide whether you actually need ongoing tracking. For most sites under 15 high-value queries, periodic baselines after content changes are enough — weekly monitoring is over-instrumentation.

Week 3. If going paid, trial two tools side by side. Check whether their reported rates match your manual baseline within a reasonable margin. A tool reporting 80% where you measured 20% is sampling favourably or has a methodology problem worth asking about.

Week 4. Commit to a tool or a manual schedule, and write down the methodology so future-you interprets the data consistently.

Past 30 days, keep it lean. The point isn't maximum data collection — it's knowing whether the work is paying off, which needs only the metrics tied to decisions you'd actually make.

AI
Free tool · No signup
Free AI Crawler Checker
Paste a domain → see which AI crawlers your robots.txt allows and which it blocks, with training and retrieval bots told apart.
Check your crawlers

FAQ

Can I see AI Overview citations in Google Search Console?
Only partially. Search Console's generative AI report, launched June 2026, shows AI Overview and AI Mode impressions broken down by page, country, device and date. It does not include clicks, click-through rate, position, or which queries cited you, and it shows nothing about the answer content or which other sources appeared. It's useful for directional trends but cannot tell you whether a specific page is cited for a specific query — that still requires manual sampling or a third-party tool.
How many times should I run a query to get a reliable citation rate?
Five to ten runs, spread across at least three days. Citation sets are unstable enough that single observations tell you nothing — SE Ranking measured 42.4% of previously cited domains being replaced when Google switched AI Overviews to Gemini 3, with no pages changing. Five runs gives a rough rate that distinguishes four-out-of-five from one-out-of-five. Ten gives a rate you can compare against a later measurement. Below five, draw no conclusions.
Is there a free way to track AI citations?
Yes, with time investment. Manual sampling in a spreadsheet works well for up to about 20 queries across two or three engines and costs only hours. The Perplexity API adds cheap automation for one engine — a small script running 20 to 50 queries weekly costs a few dollars a month. Google AI Overviews, ChatGPT Search and standalone Gemini have no clean API path, so they remain manual. Realistic free stack: Perplexity API plus manual sampling elsewhere, two to four hours weekly.
Which engine should I start tracking first?
Perplexity, because it cites on nearly every answer and therefore produces a readable citation rate within a week. ChatGPT is the hardest starting point despite its size — analysis of US desktop traffic found only 6.8% of queries carried a citation, so most of your sample rows come back empty and you learn little about your own content. Google AI Overviews sit in between: wide reach, but results vary by session and location, which makes consistent sampling harder.
How much should I budget for AI citation tracking tools?
Published entry pricing currently runs from $29 to about $89 per month, with mid-tiers around $189 and top published tiers near $489. Otterly starts at $29, AIclicks at $59, LLMrefs at $79 with a free tier, and Peec around $89. Profound does not publish self-serve pricing. Before committing, run 30 days of manual baseline tracking — most teams discover they care about fewer queries than they assumed, which changes which tier they need.
A tool says I'm cited but I can't reproduce it manually. Who's right?
Probably both, at different moments. Citation sets vary between runs, so the tool may have sampled when you appeared and you may have checked when you didn't. Resolve it by running the query eight to ten times yourself. If your rate lands within 10 to 20 percentage points of the tool's, both are accurate at the rate level. If the tool reports 80% where you measure 20%, ask the vendor how many runs per prompt it performs — that gap usually means shallow sampling rather than a bug.
How long after a content change should I expect citations to move?
Four to eight weeks. The page needs re-crawling, then re-indexing, then re-evaluating against queries, and the last step is the slow one. Perplexity tends to be faster because its index updates more frequently — changes are often visible within two to four weeks. Google AI Overviews are slower. Record a baseline before you change anything, then sample again at four weeks; comparing to an impression of how things used to be produces false conclusions in both directions.

Closing

Citation tracking is measurement without instrumentation. No engine gives publishers the data, so everyone — including every paid tool — is sampling public surfaces and calling it a rate.

That makes methodology the whole discipline. Sample enough runs that the number means something. Spread them across days. Record a baseline before you change anything. And treat any single observation, in either direction, as one data point rather than as news.

Start with 10–15 queries on Perplexity, manually, for a month. It costs nothing, it tells you which queries you actually care about, and it gives you a baseline that makes everything you do afterwards measurable.

The diagnostic for a citation that disappears is in that guide; the tool landscape in the AEO tools comparison; the underlying mechanics in what is AEO.


Sources cited in this piece

  • Google Search Central — Search Console generative AI report, launched June 2026, reporting AI Overview and AI Mode impressions only.
  • SE Ranking — Gemini 3 before-and-after analysis across 100,000 keywords; 42.4% of previously cited domains replaced; sources per answer up 31.8%, from 11.55 to 15.22.
  • Google — Gemini 3 rolled out globally as the default model for AI Overviews on 27 January 2026.
  • Profound — dataset of 680M+ AI citations; Wikipedia the top ChatGPT-cited domain at 7.8%, Reddit the top domain for Perplexity at 6.6% and Google AI Overviews at 2.2%.
  • TechCrunch (May 2026) — 6.8% of US desktop ChatGPT queries carried a citation.
  • Vendor pricing from each tool's published pricing page at the time of writing; Profound does not publish self-serve pricing.

Last updated 13 August 2026: removed the claim that 45.5% of AI citations change per re-run — no primary source could be found — and rebuilt the volatility argument on SE Ranking's measured Gemini 3 churn. Corrected the Search Console AI report description: it launched June 2026 and reports impressions only. Updated all tool pricing, which had changed across the category.

Written by
Ed Grows
Solo founder of citelity. Building AEO tools. Documenting what works (and what breaks) on aivario.com.
← Back to all posts