← Learning Hub

Learning Hub · Tideflow AI

How Do You Measure AI Citations When Google Search Console Won’t Show Them?

Search Console now shows part of the picture, but only part, so you need three separate measurements: which answers name and cite you, which AI bots discover your pages, and which AI referrals turn into conversions.

Three separate analog measuring instruments—a gauge, a small antenna, and a flow meter—arranged diagonally on a neutral tabletop under soft directional light.

15 min read

Search Console now shows part of the picture, but only part, so you need three separate measurements: which answers name and cite you, which AI bots discover your pages, and which AI referrals turn into conversions.

Google Search Console reports when links to your site appear in Google's AI Overviews and AI Mode. It does not show the prompts behind those appearances, unlinked brand mentions, or anything from ChatGPT, Perplexity, or the Gemini app. Measure AI visibility in three layers that you never merge: visibility on a fixed panel of buyer prompts, AI bot discovery in your server logs, and AI-referral visits tied to conversions. This week, lock 20 to 50 buyer prompts and record a dated baseline of Mentioned, Cited, and Rank across ChatGPT, Gemini, Perplexity, Google AI Mode, and Google AI Overview. Then set up referral and conversion tracking before you change any content.

Who this is for: marketing and content teams that need to report whether AI assistants recommend their brand and whether that brings in visits and leads. Who it isn't for: teams looking for a single number that proves AI revenue. No such number exists, and this guide explains why. You need access to Search Console, your analytics tool, and server or CDN logs. The vocabulary (mention, citation, recommendation) is defined in AI Mentions, Citations, and Recommendations: What Actually Matters?. This guide covers how to measure them.

The three measurement layers at a glance

Each layer answers a different question. A strong result in one layer tells you nothing about the others.

Scroll horizontally to see all columns →

LayerQuestion it answersWhat to trackWhat it cannot prove
1. VisibilityDo AI answers name and cite us for buyer questions?Mentioned, Cited, and Rank per prompt, per engine, over repeated runsThat anyone clicked or bought
2. DiscoveryCan AI systems find and fetch our pages?Verified AI crawler and fetcher requests in server logsThat a fetched page was cited
3. BusinessDoes AI-driven attention produce visits and revenue?AI-referral sessions and conversion eventsThe full extent of AI influence, since many visits can't be identified

Google Search Console's AI report adds one more input. It only covers Google surfaces, and it sits inside Layer 1 as a partial citation signal.

What Search Console's AI report shows, and what it misses

Search Console does report Google AI activity now, so "GSC shows nothing about AI" is out of date. Google rolled out a Generative AI performance report to all websites worldwide as of August 31, 2026. It covers impressions in AI Overviews and AI Mode, and you can break it down by date, page, country, and device (Google Search Console Help).

The limits matter more than the launch. Google defines an impression as a time "links to your site were shown to a user in a generative AI feature on Google Search." Two links from your site in the same AI feature count as a single impression in the chart total. The help page documents impressions only. It lists no clicks, no click-through rate, no position, and no query dimension. Google says it expects to update the feature list over time, so this could change.

Three practical conclusions follow:

  • GSC is a partial "Cited" signal for Google only. Because it counts link impressions, it roughly tracks citations in AI Overviews and AI Mode. If an answer names your brand without linking to you, that mention never appears.
  • You can't see which questions triggered the appearance. Without a query view, you don't know whether you showed up for "best project management software for agencies" or for your own brand name.
  • Everything outside Google Search is invisible. ChatGPT, Perplexity, and the standalone Gemini app aren't covered. Search Labs experiments are also excluded.

Each gap belongs to one of the three layers. Missing queries, mentions, and non-Google engines are covered by a prompt panel. Missing clicks and conversions are covered by referral analytics. Crawl eligibility is covered by log analysis.

Layer 1: Build a prompt panel that measures visibility honestly

A prompt panel is a fixed set of questions that you run against each AI engine on a schedule. For every answer, you record three things: whether your brand is Mentioned, whether your site is Cited as a source, and where you Rank among the brands named. It is the only way to see unlinked mentions and non-Google engines.

Write prompts in buyer language, not keyword language

Base your prompts on the questions buyers ask before they talk to sales. Sales call notes, support tickets, and comparison questions from prospects are good starting material. Most of the panel should be unbranded, option-seeking prompts, because that is where recommendations get won or lost.

Illustrative mix for a B2B scheduling tool:

  • Category discovery: "What are the best scheduling tools for small clinics?"
  • Problem framing: "How do I stop patients from missing appointments?"
  • Comparison: "Calendly alternatives for healthcare teams"
  • Branded check (a small share): "Is [Brand] HIPAA compliant?"

How much weight to give branded prompts is a judgment call, not a documented standard. Branded prompts mostly show how accurately AI describes you. Unbranded prompts show whether you get recommended at all.

Run each prompt many times, and track appearance rate instead of rank

A single AI answer is not a measurement. In SparkToro's study, 600 volunteers ran 12 prompts on ChatGPT, Claude, and Google's AI 2,961 times in total, with 60 to 100 runs per prompt. The chance of seeing the same brand list twice was under 1 in 100, and the chance of seeing the same order was roughly 1 in 1,000 (SparkToro).

Appearance rate held up across those runs. City of Hope appeared in 69 of 71 ChatGPT answers (97%) but ranked first in only 25 of them. If you tracked its rank, you'd see a brand bouncing around. If you tracked its appearance rate, you'd see a brand that is almost always recommended.

Two qualifications apply. The study was not peer-reviewed, and it was co-run with Gumshoe.ai, an AI-tracking vendor. Results were also more stable in categories with fewer candidate brands, such as cloud providers, than in crowded ones, such as sci-fi novels. The 60 to 100 runs per prompt was a research standard. No source establishes a minimum that every team must hit. We walk through a baseline-and-rerun protocol built on these findings in Are Third-Party Listicles Good GEO or Just AI-Era Link Building?.

What to do: report appearance rate per prompt and per engine, such as "Mentioned in 7 of 10 Perplexity runs." Treat rank as supporting color, not a KPI.

Split results by engine and date every baseline

Keep each engine's results separate. ChatGPT, Gemini, Perplexity, AI Mode, and AI Overview draw on different sources, and a blended score hides where you're actually weak. Put a date on every baseline and collect it again, because answers and cited sources change over time. How the answers were collected also matters. Location, logged-in state, and API versus live interface can all change the results. We cover the research on citation churn in Is AEO Really a New Discipline?.

Treat the panel as a sample you designed

A prompt panel reflects the prompts you chose, not everything buyers type. A share-of-voice figure from 50 prompts depends entirely on which 50. Keep the panel fixed between measurements, write down why each prompt is included, and add new prompts as a separate cohort so you don't quietly change your baseline.

Layer 2: Read AI bot logs as discovery, not citation

Server logs show which AI systems request your pages. A bot visit means you were discovered. It does not mean you were cited. The signal is still useful. If search-oriented bots never fetch a page, that page is unlikely to appear in those engines' answers.

The bot's identity determines what a hit means. OpenAI documents four separate agents (OpenAI):

Scroll horizontally to see all columns →

User agentWhat it doesWhat a hit tells you
GPTBotCrawls content that may be used for trainingTraining eligibility, not search visibility
OAI-SearchBotSurfaces websites in ChatGPT's search featuresYour page is discoverable for ChatGPT search
ChatGPT-UserFetches pages when a user's action triggers it. It does not crawl automatically, and robots.txt rules may not applyA live user request retrieved your page. OpenAI says this agent does not determine what appears in Search
OAI-AdsBotVisits only submitted ad landing pagesNothing about organic visibility

Verify before you count. Anyone can fake a user-agent string. Match claimed OpenAI bots against the IP range files OpenAI publishes, or use your CDN's bot verification, and label anything you can't verify as unverified. Other AI companies run their own crawlers, including PerplexityBot and Anthropic's ClaudeBot. Look up each vendor's documentation before you read meaning into their hits.

Google has a similar distinction. Blocking Google-Extended limits training of the models behind Search AI features. It does not remove you from AI Overviews or AI Mode. That is handled by a separate Search generative AI control in Search Console (Google Search Console Help).

Layer 3: Tie AI referrals to conversions

The business layer answers the question leadership actually asks: does AI visibility produce visits and revenue?

Start with ChatGPT's UTM tag. ChatGPT automatically adds utm_source=chatgpt.com to referral URLs from its search results (OpenAI Help Center). This is the most dependable way to identify AI referral traffic. In your analytics tool, create an AI-referral channel or segment that matches:

  • utm_source=chatgpt.com
  • referrer domains such as chatgpt.com, perplexity.ai, and gemini.google.com

Then connect that segment to real conversion events, such as demo requests, signups, or purchases, rather than pageviews. A thousand AI-referred sessions with no conversions is a finding worth reporting.

Accept that some AI traffic will stay hidden. The UTM tag only covers links clicked in ChatGPT search results. Copy-pasted URLs, app handoffs, and stripped referrers show up as direct traffic. A buyer who reads an AI answer and then types your brand into Google shows up as branded search. So your AI-referral count is a floor, not a total. Report it that way instead of trying to estimate the missing share.

Test whether a content change actually moved anything

When you publish or update content to close a visibility gap, you want to know whether it worked. Because AI answers vary so much from run to run, a before-and-after screenshot proves nothing. Follow this protocol:

  1. Freeze the prompts that the change targets, plus a few control prompts it shouldn't affect.
  2. Record a dated baseline by running each prompt several times on each engine and noting appearance rate for Mentioned and Cited.
  3. Make the change and confirm in your logs that search-oriented bots have fetched the new page.
  4. Rerun the same prompts the same number of times after a set interval, and compare appearance rates. Ignore rank.
  5. Check the control prompts. If they moved as much as the targets, you are probably seeing normal variation rather than your change.

Even when this works cleanly, it shows correlation, not proof. Repeated runs and controls make a real effect more believable. They don't make it certain.

Mistakes that make AI visibility reports misleading

  • Reporting "12 of 20 prompts" as pipeline. A mention rate is a visibility metric. Many AI answers produce no click, and clicks don't always convert. Report the three layers side by side, never as one chain of multiplied numbers.
  • Treating a vendor visibility score as a market metric. Any score reflects that vendor's prompt set, collection method, and number of runs. When you evaluate a tool, ask what prompts it uses, how often it runs each one, and whether it collects answers from APIs or live interfaces.
  • Reading a GPTBot spike as progress. GPTBot is a training crawler. For ChatGPT search visibility, OAI-SearchBot is the agent to watch.
  • Buying ads to "boost GEO." Nothing in OpenAI's crawler documentation links ad spend to organic citations, and its ads bot is a separate agent that only visits ad landing pages. Measure paid and organic AI visibility separately.
  • Opting out of the wrong control. Blocking Google-Extended won't remove you from AI Overviews. Turning on the Search generative AI control will, within one to two days.

How to report AI visibility to leadership

Use one short weekly or monthly view with a line for each layer, and don't add them together:

  • Visibility: appearance rate for Mentioned and Cited on the fixed panel, by engine, against the dated baseline. Include GSC AI impressions as the Google-only citation line.
  • Discovery: verified search-bot fetches of priority pages, with new pages marked once they've been picked up.
  • Business: AI-referral sessions and conversions, clearly labeled as a minimum count.

Add one sentence on what changed and one on what you'll do next. That keeps the report honest and still useful for decisions.

Where Tideflow fits in this workflow

The three layers usually live in three different tools: a prompt tracker, a log pipeline, and an analytics platform. Then someone has to connect them by hand. We built Tideflow AI to run them as one loop. We make Tideflow, so read this section with that in mind.

  • Visibility: we monitor prompts across ChatGPT, Gemini, Perplexity, Google AI Mode, and Google AI Overview. For each prompt we report Mentioned, Cited, and Rank, along with the competing brands and the sources AI relied on (Tideflow AI). We don't monitor Claude, Copilot, Grok, or DeepSeek answers, and we don't currently publish our per-prompt sampling frequency.
  • Discovery: we log AI bot requests server-side through a Cloudflare Worker or Express middleware. A visit is marked verified only when Cloudflare's bot verification confirms it, and Express observations stay unverified (Tideflow AI vs Profound). Don't run both collectors on the same traffic, because that double-counts visits.
  • Business: our analytics attribute AI referrals and conversions to the session's landing page using single-session attribution, not multi-touch. Only events you flag as conversions count as conversions. Not every AI-originated visit can be recognized (Tideflow AI developer docs).
  • Action: when monitoring shows a gap, we turn it into a content plan and publish through MCP into your own AI coding agent. The agent confirms the exact revision it applied, and deployment is tracked separately (Tideflow AI). This assumes an MCP-connected agent, so teams on a traditional CMS workflow should check how the handoff would fit.

Tideflow is in early access, and we can't guarantee any brand will be cited or recommended. Other platforms take different approaches and may be a better fit for specific jobs. We compare them directly in Tideflow AI vs Profound and Tideflow AI vs Peec AI.

Frequently Asked Questions

Why doesn't the Generative AI report appear in my Search Console property?

The Generative AI performance report can be missing if your site has too few impressions in AI features, or if the property has been excluded through the Search generative AI control. Include is the default, and child properties inherit their parent's setting, so check the parent property's setting first. Blocking Google-Extended does not cause this, because it only affects training.

How many prompts should an AI visibility panel have, and how often should they run?

Twenty to fifty buyer prompts is a practical starting point that covers your main categories without becoming unmanageable. Run frequency is less settled. SparkToro used 60 to 100 runs per prompt as a research standard, but no source sets a minimum for working teams. Stable categories with few competitors need fewer runs than crowded ones. Whatever you pick, keep it the same before and after any change you want to measure.

Can AI visibility scores from vendors be trusted?

Treat a vendor's AI visibility score as a consistent internal trend line, not an objective market share. Scores depend on the vendor's prompt set, number of runs, and collection method, whether API or live interface. Before trusting one, ask which prompts it covers, how many runs sit behind each data point, and whether you can see the individual answers.

Your next step

Start with the baseline this week: pick 20 to 50 buyer prompts, record Mentioned and Cited appearance rates on each engine, and date the result. Next, add the utm_source=chatgpt.com and referrer rules to your analytics and connect them to real conversion events. Change content only after both are in place, so that your first optimization produces a comparison you can defend.

Sources

Explore more in Learning Hub →