To measure AI search visibility, start with the one report that officially exists, add two metrics you can build with a spreadsheet, and refuse to pay for the rest until the data they claim to measure exists. That's the whole method. The industry version has more dashboards; it does not have more information.
The measurement problem is real, though. Search is shifting from links to answers, and the numbers that ran every SEO report for twenty years (rankings, impressions, click-through rate) describe a shrinking share of how customers meet you. A16z-flavored analysis of the shift, via a2aprotocol.ai, tracks average query length growing from about 4 words to about 23, with session depth stretching toward six-minute conversations. People aren't searching less; they're searching in a shape your rank tracker can't see.
What you can measure today
Three surfaces work right now, and only one of them is new.
The official one: Search Console's Generative AI performance report, Google's measurement surface for your visibility in AI features. If AI Overviews and AI Mode matter to your traffic, this is where Google itself says to look. It's also the report most audits still skip, which makes it the cheapest differentiation in reporting today.
The buildable one: citation rate. Write down the twenty to fifty questions your customers ask. Ask them in ChatGPT, Perplexity, and Google, monthly, and record who gets cited and who gets recommended. This is a spreadsheet and a repeatable habit, not a platform. Its sibling is unprompted brand recognition: does the model mention you when the question doesn't? That's the AI era's brand-awareness metric, and a monthly test tracks its direction.
The familiar one: AI referral traffic in your analytics, which we've covered in how to measure your AI traffic. Small numbers for most sites, but trend direction is the point.
One data point worth holding onto while the transition reshuffles your traffic: Google reported (May 2024) that when an AI Overview includes links, those links earn more clicks than a traditional listing for that query. The lesson isn't that everything is fine; it's that the traffic is moving, not purely vanishing, and measurement should follow where it goes.
The citation test, specified
Since citation rate is the metric you'll build yourself, here's the recipe with the corners filled in. Pick 20 to 50 questions in three tiers: buying questions ("who's the best [service] in [city]," "how much does [service] cost"), comparison questions ("[you] vs [competitor]," "alternatives to [category leader]"), and category questions where your name shouldn't appear in the prompt but should appear in the answer. Freeze the list; a rotating question set measures your curiosity, not your visibility.
Then, monthly, same week each month: ask each question in ChatGPT, Perplexity, and Google (in a clean session, not your logged-in personalized one), and log four columns per question: cited (your content referenced), recommended (your business named as an answer), competitor cited instead, or absent. The month-over-month movement in those columns is your AI visibility trend, and after a quarter you'll know which content earns citations and which competitors own which questions. Total cost: an hour a month and a spreadsheet.
What's coming: the KPIs of completed tasks
The agentic web changes the unit of success from the visit to the finished job. The forward KPI stack, from the A2A protocol layer: successful task handoffs (an agent delegated a booking, order, or quote to your systems, and it completed), agent invocation counts (how often agents call your endpoints at all), and secure data exchanges between your systems and agent ecosystems. Where clicks measured attention, these measure execution.
You should know these terms because they're the direction of travel, and because your infrastructure choices this year decide whether the events will ever exist to count. What you shouldn't do is buy a dashboard for them yet.
The candid half: what nobody can measure yet
Here's the part that shortens our own pitch. Roughly half of this KPI stack cannot currently be instrumented by anyone, for a structural reason: the agent traffic that would generate the events barely exists. Task handoffs need agents delegating tasks. Invocation counts need published endpoints for agents to invoke. Most businesses have neither, because the protocol layer is still in enterprise pilots.
So the three-way split every report should disclose looks like this: measurable now (Generative AI report, citation rate, unprompted recognition, AI referrals, plus one early agentic signal: agent visits in your server logs, identifiable since Google's dedicated Google-Agent user agent arrived in March 2026); measurable soon with preparation (invocations, handoffs, agent-referred conversions); measurable by nobody today regardless of what the sales deck implies. A vendor quoting live task-handoff analytics for a business with no agent endpoints is selling you a dashboard for weather on a planet you haven't landed on. The wider renegotiation these KPIs are tracking, and who loses in it, is mapped in the end of the Google bargain.
The worksheet
The giveaway, copy it into a doc and fill in the right column:
TRACK NOW (monthly)
[ ] GSC Generative AI performance report: impressions/clicks trend
[ ] Citation test: 20-50 customer questions across ChatGPT,
Perplexity, Google; record citations + recommendations
[ ] Unprompted recognition: category questions without your name
[ ] AI referral sessions in analytics
[ ] Google-Agent hits in server logs (count, pages requested)
INSTRUMENT NEXT (as agent readiness work proceeds)
[ ] Endpoint response times and uptime
[ ] Structured first-party data coverage (services, prices, hours)
[ ] Agent-referred conversions (tag when the events exist)
DOESN'T EXIST YET (decline dashboards for these)
[ ] Task handoff rate [ ] Invocation share vs competitors
[ ] Agent "ranking" positions
The pattern behind the whole stack: measurement follows infrastructure. The businesses whose data and endpoints are agent-ready will be the first with real numbers in the bottom section, and the readiness work is the same fundamentals Google keeps pointing at. We run this exact split (measured now, instrumented next, flagged as unmeasurable) in the Assay phase of the Reforge Method, because a report that mixes those three categories isn't a report, it's a pitch.