İçeriğe geç
SEO & GEO

Measuring AI Visibility: GEO KPIs and a Monitoring Dashboard

Are you visible in AI answers, and how would you know? GEO KPIs across four layers: crawl, citation, referral and conversion, plus how to build the dashboard.

Ahmet Berk ArslanLast Updated: 10 August 2026

If you do not know whether your GEO work is paying off, what you are doing is not optimization. It is guessing.

In our earlier GEO guide we covered the technical side of becoming visible in AI search: llms.txt, schema.org, server-side rendering, WebMCP. That article had a GEO versus SEO comparison table, and one row in it was honestly left blank:

Measurement tooling: Not yet mature - brand mention tracking, manual testing.

This article fills that row. What to measure, which metrics actually tell you something and which are just reassuring, and how to build the dashboard that collects them on your own infrastructure.

Quick Answer: AI visibility is not one metric but four distinct layers: the bot fetching your page (crawl), the model naming you in its answer (citation), the user clicking through (referral), and that traffic turning into business (conversion). Each layer is read from a different data source, and those sources are not interchangeable. Crawl data lives in server logs and is invisible in Google Analytics. Citation data has no official API and must be sampled with a fixed prompt set. Referral data shows up in analytics as a referrer. Until you bring all four together in one place, the question "is GEO working" has no answer.

Table of Contents

  • Why is this hard to measure?
  • The four layers of AI visibility
  • The GEO KPI table
  • Which metrics are trustworthy, and which are not?
  • Building the monitoring dashboard
  • A sample weekly report
  • Glossary
  • Frequently asked questions
  • Implementation checklist

Why Is This Hard to Measure?

GEO advice is abundant. There are hundreds of articles telling you to add schema, write clear definitions, publish an llms.txt. The measurement side is nearly empty, because the measurement chain that exists in classic SEO simply does not exist here.

In classic SEO the loop is closed: a query is typed, a ranking is measured, the click appears in Search Console, the traffic lands in Analytics, the conversion is written to the CRM. Every step is observable.

In AI search the middle of that chain is dark. You cannot see what the user asked ChatGPT. You cannot see which sources the model read. Nothing notifies you whether your brand appeared in the answer. And more often than not, the user gets what they needed without clicking anything at all.

So measuring AI visibility is not a matter of opening one dashboard. It is a matter of collecting four independent signals and putting them side by side.

The Four Layers of AI Visibility

Most content lumps these four together into a single bag labelled "AI visibility". But each is read from a different source, and measuring one tells you nothing about the others. Your pages might be crawled daily while you appear in no answers at all. Or you might appear constantly and still get no clicks. Those are different problems with different fixes.

1. Crawl: did the bot fetch the page?

The foundational layer. If the AI bot never visited your page, none of the other three layers can happen.

There is a critical distinction here: training crawlers and user-triggered fetches are not the same thing. GPTBot and ClaudeBot crawl broadly. ChatGPT-User and Perplexity-User, on the other hand, fetch your page because a user is asking something right now. The second is a far stronger signal: it means someone, at this moment, asked something that required your page.

2. Citation: did the model name you?

This is the one that matters, and the hardest to measure, because there is no official API for it. No AI provider gives you a dashboard saying "your brand appeared in 47 answers this week."

The only available method is sampling: you define a fixed set of questions, ask them of the models at regular intervals, and record whether your brand appeared. This will not give you absolute truth. It gives you a trend, and a trend is enough as long as you know what you are looking at.

3. Referral: did the user click?

Appearing in an AI answer does not automatically mean traffic. The user may read the answer, be satisfied, and click nothing. This is called zero-click, and its rate in AI search is far higher than in classic search.

Some people do click, and they are visible in analytics as referral traffic from sources like chatgpt.com, perplexity.ai, copilot.microsoft.com and gemini.google.com. This traffic tends to be high quality, because the user found you inside a recommendation rather than in a list of ten blue links.

4. Conversion: did it turn into business?

The final layer, and the one that determines why the other three matter. Does AI-sourced traffic fill in forms, request quotes, become revenue?

The only way to measure this is to carry first-touch source through to your form records and CRM. When a user arrives from ChatGPT and then comes back directly three days later to fill in a form, that lead is recorded as "direct" and the AI contribution disappears. Unless you store first-touch source in a cookie and submit it with the form, this layer will stay permanently empty.

The GEO KPI Table

The metrics worth tracking per layer, where to read them from, and what a healthy signal looks like:

LayerKPIData sourceHealthy signal
CrawlAI bot visits (per bot)Server / Cloudflare logsSteady and rising, spiking after new content
CrawlShare of user-triggered fetchesServer logs (ChatGPT-User, Perplexity-User)Growing as a share of total bot traffic
CrawlUnique URLs crawledServer logsApproaching your total published page count
CitationMention rate (hits per N runs)Prompt panelRising over time, comparable quarter to quarter
CitationShare of voice versus competitorsPrompt panelRising relative to competitor mentions
CitationNumber of URLs citedPrompt panel (links in the answer)Spread across several pages, not just one
ReferralAI-sourced sessionsGA4 referralGrowing as a share of total organic
ReferralSession quality of AI trafficGA4 (duration, pages/session)No worse than classic organic
ReferralGen-AI impressionsSearch Console Gen-AI reportRising (no click data, see below)
ConversionAI-sourced form submissionsForm records (first-touch source)Above zero and growing
ConversionCost per AI-sourced leadCRM + content production costComparable to your other channels

Which Metrics Are Trustworthy, and Which Are Not?

There is a lot of confident writing in this space. Knowing where the numbers come from is what stops you making bad calls when you look at your own dashboard.

Search Console now separates AI impressions, but gives no clicks. On 3 June 2026 Google added a dedicated Gen-AI performance report to Search Console. AI Overviews and AI Mode impressions are no longer buried inside classic organic data. However, the report gives impressions, pages, countries and devices; there is no click data, no click-through rate and no query breakdown. So the question "what is our click-through rate on the AI side" has no answer today. Google has said it will add metrics over time, but that is the current reality.

Google Analytics does not show crawlers. GA4 automatically filters known bot traffic. If you try to read the crawl layer out of GA4 the result comes back silently as zero, and there is a real risk of mistaking that for a finding. The only reliable source for crawl data is server or edge logs.

LLM answers are not deterministic. Ask the same question twice and you can get two different answers. A single prompt test is therefore an anecdote, not data. If you are measuring, you need a fixed prompt set, multiple runs per prompt, and a rate expressed as "appeared in N of M runs". A yes/no answer does not count as measurement.

Third-party GEO tools cannot see inside the model. The AI visibility tools on the market use exactly the same sampling method: they ask their own prompt sets and count the results. The share-of-voice number they hand you is an estimate of that set, not absolute truth. Using a tool can be sensible. Treating its output as ground truth is not.

Two numbers for scale. AI search visits grew 42.8% year over year between Q1 2025 and Q1 2026, from 15.6 billion to 27.4 billion. Against that, only 14% of marketers track AI citations, while 43% name AI search optimization a core strategy for 2026. That is close to a threefold gap between the people calling it a strategy and the people measuring it. Starting to measure is still, today, a point of differentiation.

Building the Monitoring Dashboard

The setup below assumes Next.js, MongoDB, Cloudflare and GA4. Adapting it to your own stack is not difficult; the logic is the same.

Collecting crawl logs

You isolate AI bots by user-agent. Counting training crawlers separately from user-triggered fetches matters, because the latter is a much stronger intent signal:

# Training and search crawlers
CRAWLERS='GPTBot|ClaudeBot|Google-Extended|Bytespider|CCBot|Applebot-Extended'

# User-triggered fetches - someone is asking something right now
USER_FETCH='ChatGPT-User|Perplexity-User|OAI-SearchBot|PerplexityBot'

# Which bot fetched which page, how many times
grep -E "$CRAWLERS|$USER_FETCH" access.log \
  | awk -F'"' '{print $6, $2}' \
  | sort | uniq -c | sort -rn | head -30

If you are on Cloudflare you can push the same data to storage via Logpush and aggregate it daily. What matters is accumulating the number somewhere, because a one-off look tells you nothing. You need the trend.

Defining the prompt panel

Citation measurement rests on a fixed question set. Fixed is non-negotiable: if the questions change between measurements, you cannot compare anything. Derive the questions from real customer questions, not from questions that already contain your brand name:

{
  "brand": "Detartech",
  "competitors": ["Competitor A", "Competitor B", "Competitor C"],
  "locale": "en-US",
  "runsPerPrompt": 5,
  "schedule": "0 6 * * 1",
  "prompts": [
    {
      "id": "geo-agency-en",
      "text": "Can you recommend a software agency that does GEO and AI search optimization?",
      "intent": "commercial"
    },
    {
      "id": "mvp-agency-en",
      "text": "Which software companies build MVPs in Istanbul?",
      "intent": "commercial"
    },
    {
      "id": "llms-txt-en",
      "text": "Does llms.txt actually work?",
      "intent": "informational"
    }
  ]
}

Three details matter. runsPerPrompt must be greater than one, because a single run is an anecdote. The intent field lets you separate commercial from informational questions; appearing in a commercial answer is worth far more. And without a competitors list you cannot compute share of voice at all.

Running the panel

A simple loop that asks the prompt set and records the result. The example below uses a model with web search enabled, because what you want to measure is not the model's training data but the answer it produces from the live web:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

type PromptRun = {
  promptId: string;
  runIndex: number;
  mentioned: boolean;
  competitorsMentioned: string[];
  citedUrls: string[];
  answer: string;
  checkedAt: Date;
};

async function runPrompt(
  prompt: { id: string; text: string },
  brand: string,
  competitors: string[],
  runIndex: number,
): Promise<PromptRun> {
  const response = await client.messages.create({
    model: 'claude-opus-5',
    max_tokens: 4096,
    tools: [{ type: 'web_search_20260209', name: 'web_search' }],
    messages: [{ role: 'user', content: prompt.text }],
  });

  const answer = response.content
    .filter((block) => block.type === 'text')
    .map((block) => block.text)
    .join('\n');

  return {
    promptId: prompt.id,
    runIndex,
    mentioned: new RegExp(brand, 'i').test(answer),
    competitorsMentioned: competitors.filter((c) =>
      new RegExp(c, 'i').test(answer),
    ),
    citedUrls: [...new Set(answer.match(/https?:\/\/[^\s)\]]+/g) ?? [])],
    answer,
    checkedAt: new Date(),
  };
}

Store the raw answers too. When your mention rate drops, the explanation is in those texts: who the model started recommending instead of you, which source it cited. You can only see that by reading the answer.

Separating referral traffic

You classify AI sources by referrer and persist the result as first-touch source. The conversion layer depends entirely on this:

const AI_REFERRERS: Record<string, string> = {
  'chatgpt.com': 'ChatGPT',
  'chat.openai.com': 'ChatGPT',
  'perplexity.ai': 'Perplexity',
  'copilot.microsoft.com': 'Copilot',
  'gemini.google.com': 'Gemini',
  'claude.ai': 'Claude',
};

export function classifyReferrer(referrer: string | null): string | null {
  if (!referrer) return null;
  try {
    const host = new URL(referrer).hostname.replace(/^www\./, '');
    return AI_REFERRERS[host] ?? null;
  } catch {
    return null;
  }
}

Write this value to a cookie on first visit and attach it to the form payload on submission. Otherwise the user who returns directly three days later gets recorded as "direct" and the AI contribution becomes invisible.

Bringing the data together

Collect all three sources in one collection and produce a weekly summary:

// geo_metrics collection: { layer, metric, value, source, recordedAt }

db.geo_metrics.aggregate([
  { $match: { recordedAt: { $gte: new Date(Date.now() - 7 * 864e5) } } },
  {
    $group: {
      _id: { layer: '$layer', metric: '$metric' },
      total: { $sum: '$value' },
      samples: { $sum: 1 },
    },
  },
  { $sort: { '_id.layer': 1, total: -1 } },
]);

A Sample Weekly Report

Once the dashboard is running, a weekly read looks something like this:

LayerMetricThis weekLast weekReading
CrawlAI bot visits1,2401,180Stable, nothing to act on
CrawlUser-triggered fetches9661Up, real question volume is rising
CitationMention rate40% (2/5)20% (1/5)Possibly the new content, watch two more weeks
ReferralAI-sourced sessions8377Mention gain has not reached clicks yet
ConversionAI-sourced forms32Low volume, review monthly not weekly

The real skill here is not reacting to a single week's movement. Mention rate is sampled, so it fluctuates by nature. Look at a four-week trend before making a decision.

Glossary

  • GEO (Generative Engine Optimization): The practice of structuring content so generative AI models can use it as a source in their answers.
  • Mention rate: The proportion of runs in a fixed prompt set where the brand appeared.
  • Share of voice: Your portion of all brand mentions across those same answers.
  • Prompt panel: The fixed set of questions asked of models at regular intervals to measure citation.
  • Zero-click: The user getting their answer and ending the search without clicking any site.
  • User-triggered fetch: A page fetch caused by a user asking something right now (ChatGPT-User, Perplexity-User).
  • First-touch source: The channel where the user first encountered the site, stored so it is not lost on later visits.
  • Gen-AI performance report: The Search Console report that separates AI Overviews and AI Mode impressions.

Frequently Asked Questions

Can I measure AI visibility with a single metric? No. Crawl, citation, referral and conversion are separate layers and one tells you nothing about the others. You can be crawled daily and appear in no answers at all.

Does Search Console show my AI clicks? No. The Gen-AI report released in June 2026 gives impressions, pages, countries and devices; it does not give clicks, click-through rate or query breakdown.

How many questions should the prompt panel have? 15 to 30 is enough for most businesses. What matters is not the count but that the questions derive from real customer questions and stay fixed over time. Change the set and you lose comparability.

Should I buy a GEO tool or build this myself? Either works. Off-the-shelf tools start fast, but their sampling is not specific to you and they will not wire the conversion layer into your CRM. Your own dashboard is more work, but it fills the crawl and conversion layers with data that is genuinely yours.

How soon will I see results? The crawl layer responds within days. Citation is much slower and much noisier; you need at least four weeks of data for a meaningful trend. The conversion layer is low volume and should be reviewed monthly.

My mention rate dropped. What now? Read the raw answers first. Who did the model start recommending instead of you, and which source did it cite? More often than not the problem is not on your page but that a competitor published a clearer, more quotable definition.

Implementation Checklist

  • Report AI bot user-agents separately in your server or edge logs
  • Count training crawlers and user-triggered fetches separately
  • Define a fixed prompt set of 15 to 30 questions and do not change it
  • Run each prompt at least 3 to 5 times, never once
  • Store raw answers, not just a yes/no flag
  • Define a competitor list; you cannot compute share of voice without one
  • Set up AI referral sources as a dedicated segment in GA4
  • Persist first-touch source in a cookie and attach it to form submissions
  • Check the Search Console Gen-AI report weekly
  • Look at a four-week trend before deciding anything

Closing

Measuring AI visibility is not as mature as measuring classic SEO today, and it will not be for a while. But "cannot be measured" and "does not arrive in a single dashboard" are two different statements.

Collect the four layers separately and put them side by side, and you end up with enough data to support a decision: which content is being crawled, which questions you appear in, which answers produce clicks, which traffic becomes business. This is not perfect measurement. It is a great deal better than guessing.

We run the technical GEO work under our GEO service and the classic search side under our SEO service. For the technical side of becoming visible, see our GEO guide; for how AI agents will actually operate websites, see our WebMCP article.

Have a project in mind?

Let's bring the technologies from this article to life in your project.

Request a Free Discovery Call