How to track AI citations in 2026: methods that actually work
There's no dashboard for AI citations, and citation sets churn heavily. Here are the three real methods, what to measure, and which numbers are noise.
Tracking AI citations is manual work, and the unit of measurement is a rate rather than an event. No engine offers a Search Console equivalent, citation sets vary considerably between runs, and each engine exposes different amounts of data.
Three approaches exist: manual sampling (free, good to about 20 queries), semi-automated scripting (cheap, moderate setup), and paid tracking tools ($29–$500/month). This guide covers what each does, what to measure, and which numbers are signal.
Disclosure: I'm building citelity, which includes citation tracking. I've been specific below about what it does and where other tools fit better.
Why is AI citation tracking hard?
Three structural problems, and they compound.
No data is exposed. Google Search Console added a generative AI report in June 2026 showing AI Overview and AI Mode impressions — broken down by page, country, device and date. It contains no clicks, no CTR, no position, and no query-level citation data. For ChatGPT, Perplexity, Claude and standalone Gemini there is no publisher dashboard at all.
Citation sets churn. SE Ranking analysed 100,000 keywords before and after Google switched AI Overviews to Gemini 3 on 27 January 2026: 42.4% of previously cited domains were replaced and sources per answer rose 31.8%, from 11.55 to 15.22. No page changed — the model did. If a model swap can move that much, a single observation proves nothing.
Engines differ. Profound's dataset of over 680 million citations shows distinct source preferences per engine — Wikipedia leads ChatGPT citations at 7.8%, Reddit leads both Perplexity at 6.6% and Google AI Overviews at 2.2%. A tool that samples one engine well often handles the others poorly.
What should you measure?
Five metrics, in order of how much they tell you.
What's mostly noise: a single citation appearance; day-of-week patterns below 50 samples; position changes of one or two places; cross-engine comparisons on queries where one engine doesn't generate an AI answer at all.
Method 1: manual sampling
The right starting point for almost everyone. Free, scales to roughly 20 queries, and it teaches you what to look for before you spend anything.
Setup. A spreadsheet with columns for query, date, engine, session type (regular or incognito), cited yes/no, position if cited, and the extracted sentence if identifiable.
Routine. For each query: run it five to ten times across at least three different days, mixing regular and incognito sessions since personalisation can inflate appearance. Record each run as a row. After eight to ten runs, calculate the rate.
The arithmetic. Fifteen queries across four engines at ten runs each is 600 samples. At 30–60 seconds per run that's five to ten hours — feasible as a one-time baseline, not as weekly monitoring.
Where it stops working: past 20 queries, past two or three engines, or when you need to detect change within a week of an edit.
Method 2: semi-automated
Between fully manual and paid, there's a band that automates part of the work.
Perplexity API is the cleanest option. It returns citation data structurally, so a small script can run 20–50 queries weekly and log results to a database for a few dollars a month. It covers one engine well and nothing else.
Screenshot workflows — a full-page capture extension plus a rotating daily schedule — get you to roughly 25–35 queries weekly at about 30 minutes a day, with better records than pure manual.
Scraping scripts for AI Overviews exist on GitHub and break regularly as Google adjusts rendering. Budget for maintenance, or don't start.
ChatGPT Search and standalone Gemini have no clean API path for search-grounded responses. Treat both as manual-only surfaces.
A realistic stack: Perplexity API scripted weekly, manual sampling for AI Overviews and ChatGPT, spreadsheet aggregation. Roughly $5–20/month and two to four hours weekly.
Method 3: paid tracking tools
Current published pricing, which has moved considerably in this category over the past year.
| Tool | Entry | Notes |
|---|---|---|
| Otterly.AI | $29/mo | Standard $189, Premium $489. Cheapest genuine on-ramp |
| AIclicks | $59/mo | Pro $189, Business $499. Attribution-focused |
| LLMrefs | $79/mo | ~500 prompts, up to 50 keywords, plus a free tier |
| Peec AI | ~$89/mo | Tiered by prompt volume; strong competitive analysis |
| Profound | not published | Demo-led, enterprise custom. Largest public dataset |
What to ask before buying:
Sampling depth. How many runs per prompt does the tool actually perform? A tool reporting citation status from one or two samples is reporting a coin flip. Five or more is the minimum for a meaningful rate. Few vendors publish this — ask.
Engine coverage. Does it cover the engines your audience uses? AI Overviews plus one other is the minimum useful pair for most content sites.
Snippet capture. Does it record which sentence was quoted? This is the most actionable detail and tools vary widely on whether they capture it.
Cost per tracked prompt. Divide monthly cost by prompt slots. A $189 plan covering 350 prompts is cheaper per unit than a $79 plan covering 50.
citelity (mine). Tracking sits inside a Find → Fix → Prove loop rather than standing alone: it starts from Search Console to find pages losing clicks while rankings hold, generates the fix, then measures the result in positions, citations and GA4 traffic. Tracked engines are ChatGPT, Perplexity, Claude and Google AI Overviews. $49 / $99 / $199 a month. If you only need tracking and already have a content process, a visibility-only tool is cheaper and I'd point you there.
How often should you sample?
Five mistakes that produce wrong conclusions
Claiming a win from one observation. You got cited this morning. That's one sample. Don't celebrate — or post about it — until you have a rate from five or more runs.
Comparing engines on incompatible queries. "Cited in Perplexity but not AI Overviews" might be a structural pattern, or the query might not trigger an AI Overview at all. Check that both engines actually generate an AI answer before drawing a conclusion.
Reading micro-position changes. Moving from cited source four to source three is within variance. Presence versus absence, and the rate, are the actionable measurements.
Sampling all at once. Ten queries in one hour gives you one session, not a pattern over time. Spread across at least three days.
Comparing against memory. "It feels like we were cited more last month" is not a baseline. If you didn't record it, you can't compare to it — and this is the single most common reason people conclude their AEO work isn't paying off.
A realistic 30-day setup
Week 1. Pick 10–15 queries you genuinely care about. Run ten samples each on AI Overviews plus one other engine. Record everything. Calculate baseline rates.
Week 2. Decide whether you actually need ongoing tracking. For most sites under 15 high-value queries, periodic baselines after content changes are enough — weekly monitoring is over-instrumentation.
Week 3. If going paid, trial two tools side by side. Check whether their reported rates match your manual baseline within a reasonable margin. A tool reporting 80% where you measured 20% is sampling favourably or has a methodology problem worth asking about.
Week 4. Commit to a tool or a manual schedule, and write down the methodology so future-you interprets the data consistently.
Past 30 days, keep it lean. The point isn't maximum data collection — it's knowing whether the work is paying off, which needs only the metrics tied to decisions you'd actually make.
FAQ
Can I see AI Overview citations in Google Search Console?
How many times should I run a query to get a reliable citation rate?
Is there a free way to track AI citations?
Which engine should I start tracking first?
How much should I budget for AI citation tracking tools?
A tool says I'm cited but I can't reproduce it manually. Who's right?
How long after a content change should I expect citations to move?
Closing
Citation tracking is measurement without instrumentation. No engine gives publishers the data, so everyone — including every paid tool — is sampling public surfaces and calling it a rate.
That makes methodology the whole discipline. Sample enough runs that the number means something. Spread them across days. Record a baseline before you change anything. And treat any single observation, in either direction, as one data point rather than as news.
Start with 10–15 queries on Perplexity, manually, for a month. It costs nothing, it tells you which queries you actually care about, and it gives you a baseline that makes everything you do afterwards measurable.
The diagnostic for a citation that disappears is in that guide; the tool landscape in the AEO tools comparison; the underlying mechanics in what is AEO.
Sources cited in this piece
- Google Search Central — Search Console generative AI report, launched June 2026, reporting AI Overview and AI Mode impressions only.
- SE Ranking — Gemini 3 before-and-after analysis across 100,000 keywords; 42.4% of previously cited domains replaced; sources per answer up 31.8%, from 11.55 to 15.22.
- Google — Gemini 3 rolled out globally as the default model for AI Overviews on 27 January 2026.
- Profound — dataset of 680M+ AI citations; Wikipedia the top ChatGPT-cited domain at 7.8%, Reddit the top domain for Perplexity at 6.6% and Google AI Overviews at 2.2%.
- TechCrunch (May 2026) — 6.8% of US desktop ChatGPT queries carried a citation.
- Vendor pricing from each tool's published pricing page at the time of writing; Profound does not publish self-serve pricing.
Last updated 13 August 2026: removed the claim that 45.5% of AI citations change per re-run — no primary source could be found — and rebuilt the volatility argument on SE Ranking's measured Gemini 3 churn. Corrected the Search Console AI report description: it launched June 2026 and reports impressions only. Updated all tool pricing, which had changed across the category.