Guide

How to measure AI visibility

Write the questions your buyers ask, run them against each engine, and record whether you were named, linked, or neither. How often to repeat depends on scale: a single prompt moved so much across 2,961 runs that one reading of it is a draw, while at 753 prompts once a day landed within two points of ten times a day.

By the Addition team Updated 10 September 2026 9 min read

Why the published procedures disagree

What to count is settled: mentions, citations, share against competitors, referral traffic. The one-off version of the same work is sold as an AI visibility audit. What is not settled is the step that decides whether the number means anything, which is how many prompts you run.

Three of them never say how many times to ask.

Three published methods never say how often to run. Semrush at least fixes what you ask, committing to a prompt set at the start of a report, and leaves the interval unstated.

Only one procedure here puts a number on both: twenty to thirty questions, fifty or more per topic, and a stated re-run interval. The rest of the readable set skips the question.

One agency states the trade in one line: 25 prompts sampled once are noisier than 500 sampled weekly. That is the whole argument for repetition, and it costs nothing to act on.

The five published methods, read 4 September 2026; the airOps guide re-read on its own site 10 September 2026Whether the procedure you are following includes a repeat count

What each published method tells you to do, side by side.

Whether the procedure you are following includes a repeat count

Practitioners still ask each other how they do this, which tells you the published metric lists are not settling it. What can be settled is smaller than another metric list, and more useful.

Why there are two right answers

The disagreement about repetition is not carelessness. Two vendors have measured it, both published their samples, and they reach opposite conclusions because they are reading different units.

Two units, two answers.

SparkToro (28 January 2026) and Profound (8 July 2026), both read 4 September 2026Which of the two your own setup is closer to

Two vendor experiments, each with its sample published, pointed at the same question from opposite ends.

Which of the two your own setup is closer to

The left-hand column has the blunter consequence: if your evidence is a screenshot of one answer, you are holding a draw.

The right-hand column has the explanation worth carrying: at that size the set is already doing the averaging. Profound also states the ceiling, which is that the platforms change underneath you, and no amount of same-day repetition beats that.

So how often to repeat turns into a question back: how many prompts are you tracking? At a handful, the set is averaging over a handful, which is not much; repeating is the other thing that would show you how far the answer moves. Neither claim rests on a test at that size, because no published one sits there. At 753 prompts, the only portfolio size anyone has tested this at, the extra runs moved the reading about two points and improved precision by roughly 10 percent. Set that against the single-prompt case. Nobody has measured where in between the effect starts. Treating 753 as a threshold, when it is only the size it was shown at, would be inventing one. Inspeccia measured that swing across 2,021 domains and found it large. That is the reason repetition belongs in the procedure at all. What the score is made of once you have it, and what that swing looks like, is the AI visibility guide.

How to measure it by hand for free

A manual version costs nothing, and Google’s own answer spells it out in five steps. Doing it by hand once tells you more than reading another metric list.

Ask buyer questions, not your own name.

Google, Profound and Clovion, read 4 September 2026Which step your current process is missing

Five steps. Four come from published sources; the fourth comes from the instability those sources measured and then left out of their own procedures.

Which step your current process is missing

The third step has a number behind it. Clovion measured that a single follow-up question changed 62 percent of the brand recommendations. So a prompt set made only of opening questions measures an opening, not a conversation. The fourth matters because named, linked and absent are separate events, and the distinction itself is in what AI visibility is.

Running that by hand, with the fourth step included, produces something a dashboard does not hand over: the spread, seen directly. After that, a tool buys you scale and history. It does not buy a number you could not have produced yourself.

What your own reporting will and will not tell you

Read what Google documents about its own reporting before you buy anything. It rules out the first plan most teams reach for, which is to find the AI number inside the reporting they already have.

Google folds these visits in. It does not break them out.

What the free report gives you, metric by metric

Google Search Console Help page for the Performance report on Search results, listing the metrics available: clicks, impressions, click-through rate and average position
  1. 1Four metrics, and Google defines each one on this page. Everything you can measure for free begins here.
  2. 2The default view is the last three months, which is the window you inherit unless you change it.
  3. 3Nothing in the metric list separates an AI feature from an ordinary result. That separation is the thing this page keeps saying you cannot get here.
Google Search Console Help, read 5 September 2026.

What Search Console means by position

Google Search Console documentation defining average position as the average of the topmost result from your site
  1. 1The figure averages the TOPMOST result from your site, not every result you have.
  2. 2So a page can lose a second listing and the reported position improves, which is the opposite of what happened.
Google Search Console Help, read 7 September 2026.

That leaves a gap between two reports: the tracker tells you whether you appeared, your analytics tells you what arrived, and neither carries the other’s unit. An appearance that produces no click is invisible to one and central to the other. Treat them as two separate reports, not one funnel. AI Overviews covers what Google has published about the surfaces themselves.

What the answer above does to the invoice

If your own reporting cannot separate these visits, you choose between running the manual procedure above and paying for scale. The scale question from two sections ago decides what the payment buys: a plan sells prompts and runs, and only one has been tested at any size.

One lever has been measured against itself. The other has not been measured at all.

The test we found varied one thing and held the other still: 753 prompts throughout, one run a day against ten. So it speaks to run frequency and is silent on prompt count. If a plan prices by runs or credits consumed per answer, the same test says the tenth run improved precision by about a tenth at that size. Neither statement has been shown at a smaller set.

That is not an argument for the cheapest plan. It is an argument for knowing which of the two you are buying, and for asking a vendor to state its own cadence before comparing its number to anything. What plans charge, in each vendor’s own unit, is set out in the AI visibility guide, and which vendors publish a cadence at all is compared in AI visibility tools.

Suppose you want the work and not the plan: measuring your own baseline is our AI visibility service.

Four ways the measurement breaks

Each of these produces a number that looks reportable. Three of them come from the same confusion about scale, and the fourth from expecting a platform to report something it has documented that it does not.

A number can be precise and still be a draw.

The mistakeWhat the evidence saysWhat to do instead
Screenshotting one answer as proofThe same list came back twice under one time in a hundred across 2,961 runsAsk it enough times to see the spread, or track enough prompts that the spread averages
Paying for runs on a set of this sizeTen runs a day improved precision by about 10 percent against one, measured at 753 promptsAt that size the extra runs bought little; whether more prompts would buy more is untested
Reporting a small prompt set as a single figureThe portfolio finding was measured at 753 prompts and nothing was tested below itReport the spread you saw, or widen the set. Where the effect begins is unpublished
Expecting AI traffic to be reported on its ownGoogle documents these appearances as counted inside the Web search type, and the documented appearance list is web, image, video, newsRead the total, and keep the tracker as a separate report
The first three are the same mistake seen from three sides, which is applying advice from one scale at the other.

A fifth sits outside your control. Both experiments behind the scale argument are published by companies selling tracking, and so are most of the methods. Their samples are stated and their methods written down. Neither makes them independent. Nobody without a product in the category has compared two trackers on the same brand, and that is the study that would make all of this checkable.

What to do this week

The two experiments behind this are vendor-published, so read them narrowly. At a handful of prompts nothing was tested, so treat your number as a small sample and let repetition show you its spread. At 753, the one size we found tested, ten runs a day instead of one moved the reading about two points and improved precision by roughly 10 percent.

Start by hand, and know which lever the evidence covers before you buy either.

Pick a small set of buyer questions. Twenty is a workable afternoon. Ask each of them more than once on the engines you care about, and record named, linked or absent for every run. Neither the number nor the repeat count is drawn from a study; what the studies establish is that a single run of a single question is not a fact. No repeat count is published for a set this small, so the useful test is not a number but a question you put to your own set: did the answer change? If it did, you have learned that one reading was never a fact, which is the thing a single screenshot hides. After that the tool buys you scale and history, and both are real. It does not buy a number you could not have produced yourself.

Sources

  1. Google Search Central AI features and your website: appearances in AI features are included in overall search traffic within the Web search type last updated 10 December 2025, accessed 4 September 2026
  2. Google Search Console Help Performance report: an impression is how many times your site appeared in Search results, and the appearance types are web, image, video and news accessed 4 September 2026
  3. Google Search Central Get started with Search Console: the report list, which does not mention AI features last updated 10 December 2025, accessed 4 September 2026
  4. Google AI Overview for “how to measure ai visibility”: a free manual method in five steps pull 2 September 2026, read 4 September 2026
  5. Brainlabs AI visibility measurement metrics: mention volume, share of voice, AI Overviews tracking, mention quality, referral traffic no date on the page, accessed 4 September 2026 This page carries no publication date of its own.
  6. Semrush How to measure AI search visibility: commit to a fixed prompt set at the start of a reporting period, then measure against it consistently. Semrush sells a visibility toolkit 20 May 2026, accessed 4 September 2026
  7. Senso.ai How to measure AI visibility: mentions, share of voice, sentiment framing and citation sources across real prompts, over time. Senso.ai sells a visibility product 22 February 2026, accessed 4 September 2026
  8. Clique Studios Measuring AI search visibility: prompt-sampled scores move run to run, so the platform must show variance, not just a single number 24 July 2026, accessed 4 September 2026
  9. airOps How to Measure AI Search Visibility: Step-by-Step Guide for 2026, 20 to 30 questions, 50 or more per topic, re-run every seven days read on airOps’ own site 10 September 2026
  10. Profound Is once a day enough: two identical setups run side by side for two weeks, 753 prompts across seven platforms, one run a day against ten. Profound sells a visibility platform 8 July 2026, accessed 4 September 2026
  11. Profound How to design prompts for AI visibility tracking: start at a hundred prompts, typical tracking runs from a hundred to a thousand, weighted towards unbranded questions 10 February 2026, accessed 4 September 2026
  12. Rand Fishkin, SparkToro AIs are highly inconsistent when recommending brands or products: 12 prompts run a combined 2,961 times. SparkToro sells audience research software 28 January 2026, accessed 3 September 2026
  13. Inspeccia Why AI visibility tools disagree: production data across 2,324 audits and 2,021 domains. Inspeccia sells a visibility product 28 July 2026, accessed 3 September 2026
  14. Clovion, reported by Greg Jarboe in Search Engine Journal Clovion data: a single follow-up question changed 62 percent of AI brand recommendations 7 July 2026, accessed 3 September 2026
  15. Google Search results for “how to measure ai visibility” read 4 September 2026

Questions people ask

How to track AI visibility?

Fix a set of buyer questions, run them against the engines you care about on a stated cadence, and record for each answer whether you were named, linked, or absent. Google publishes a free manual version of exactly that.

The part everyone leaves out is the size of the set. At a handful of prompts nothing has been tested, so the safe reading is that you are looking at single draws. At 753 prompts, the one size anyone has tested, once a day was within about two points of ten times a day and roughly 10 percent less precise.

How many times should I run the same prompt?

It depends on how many prompts you track, which is why the published answers disagree. On a single prompt, repetition is everything: 12 prompts run a combined 2,961 times returned the same brand list under one time in a hundred.

At one portfolio size it has been tested once. Running 753 prompts once a day landed within about two percentage points of running them ten times a day, and the vendor reports the extra runs improving precision by about 10 percent. That is one experiment on one platform set, not a settled rule.

Can I measure AI visibility in Search Console?

Not as a separate line. Google documents that appearances in AI features are counted inside overall search traffic under the Web search type, and the performance report’s appearance types are web, image, video and news.

So your total already includes those visits and will not isolate them. A tracker answers a different question, which is whether you appeared at all, including in the answers nobody clicked.

What is a good number of prompts to track?

One vendor’s prompt-design guide suggests starting at a hundred, and says typical tracking runs from a hundred to a thousand. It weights heavily towards unbranded questions, not your own brand name.

That is advice about coverage. The separate finding was measured at 753 prompts: a large set averages out the run-to-run movement that makes a single prompt unreadable. So the two numbers answer different questions, and neither one marks a threshold.