Guide

AI visibility audit

An AI visibility audit is a one-off report on whether answer engines name your brand. What is inside it varies. Of the audits published openly, one is entirely measurement and another is mostly diagnosis. None of the six states how many prompts it runs, which matters because a single reading of this moves between runs.

By the Addition team Updated 10 September 2026 10 min read

What an AI visibility audit includes

An audit implies a fixed scope: somebody looks at a defined list and reports what they found. Here the list is not fixed. What you get ranges from a report on where you appear to a diagnosis of why you do not.

Despina Gavoyannis at Ahrefs sets out eight steps and not one of them is diagnostic. Another audit is mostly diagnosis.

An ecommerce platform splits its audit four ways: query mapping, manual prompting, tracing each cited source back to where it came from, then an automated technical scan. That third stage is what tells you why you are missing, not just that you are.

The six audits published openly, read 4 September 2026Which of the two deliverables you are being quoted for

What each audit covers.

Which of the two deliverables you are being quoted for

One provider is explicit about which half it sells. Its page opens by calling the audit a diagnostic. It lists structural and content gaps preventing citations alongside the presence check, and that is the sentence Google’s own answer quotes when it defines the term. Not every audit sold under that name contains a diagnostic step at all, so the definition Google is repeating fits some of them and misdescribes the rest. The gap is not a quibble about labels, because the two deliverables answer different questions and cost different amounts of work. The measuring half is the same work as how to measure AI visibility, run once instead of continuously. The diagnosing half asks why you are absent. That is a question about your pages, not about the engines. That does not make its findings permanent: what counts as well-structured or as fresh is a moving standard, and one of the published checks asks exactly that.

What one reading can carry

The diagnosing half looks at your own pages, so re-running it does not depend on what an engine answered today. The measuring half is a single reading of something that moves, and that is where the missing sample size in all six starts to matter.

Three separate measurements of how far one reading can move.

The first two are published by vendors in this category and state their samples. The third reaches us through a trade publication and gives no run count.
How much a single reading of this can move

There is a condition under which one reading is enough, and it is the strongest defence of the format. Profound ran 753 prompts once a day against ten times a day for two weeks. The two readings landed within about two percentage points, because a set that size averages the movement out. So a one-off audit over a large enough prompt set is a reasonable instrument. Whether any particular audit meets that condition cannot be checked, because none of the six says how many prompts it used.

How long an audit takes: thirty days or five

The clearest way to see the tension in the format is to put two sentences from this category side by side. Both are written by companies selling into it, and they are about the same measurement.

One is an observation window. The other is a delivery promise. They are not the same kind of number, which is exactly why they are worth putting together.

SE Ranking (23 April 2026) and Ariad Partners, read 4 September 2026How long the number under your audit was observed

Two vendor statements about the same kind of measurement.

How long the number under your audit was observed

This does not make the five-day audit dishonest. A diagnosis of your pages asks whether your product information is structured, whether your expertise claims are verifiable, and whether the content matches how people ask. None of that needs thirty days of observation. It is the measuring half that inherits the thirty-day problem, and the audits that are entirely measurement inherit all of it.

Running that diagnosis on your pages, and then doing the work it names, is our AI visibility service.

Four questions to ask before buying one

The useful version of this purchase is specific about which half you are buying and about the sample behind the half that has one. Four questions cover it.

Buy the diagnosis for the fixes, and the measurement for the baseline.

Ask before you commissionWhy it changes what you getWhat a non-answer means
Is this a report or a diagnosisTwo audits contain no diagnostic step, and one is mostly diagnosisYou may receive a scorecard when you wanted a list of fixes
How many prompts, and which enginesThe count is what decides whether a single reading averages out; the engine list decides which answers were sampled at allThe number cannot be compared to anything, including your own next one
Are follow-up questions includedOne follow-up changed 62 percent of brand recommendations in one measurementThe audit measured openings, not conversations
What happens after the reportOne agency states plainly that no fixes are implemented in this phaseThe roadmap may be the deliverable, not the work
The first two are the ones every audit leaves open. The third and fourth are answered by one source each.

The third question is the one nobody asks, and it comes from a measurement. If the prompt set is built entirely from opening questions, the audit has measured how engines answer people who have not asked anything yet. One vendor’s data puts a size on the difference, with a single follow-up question changing 62 percent of the brand recommendations. Whether your own buyers ask a second question is a fact about them, and this measurement does not establish it. What it does establish is that a set of opening questions and a set including follow-ups return substantially different brand lists. On the diagnosing half, the most concrete list published anywhere is Search Engine Journal’s. Whether expertise claims are verifiable. Whether content matches how people query engines. Whether product information carries machine-readable markup, and whether content that looks fresh to you is fresh to an engine. Those are checks on your own pages, not on an engine’s output, so re-running them tomorrow does not depend on what an engine happened to answer. Whether the answer stays useful is a different question: what counts as fresh, or as well-structured, is a judgment that can change.

Search Engine Journal, read 4 September 2026Which of these your own pages would fail today

The questions this audit asks. The last one is measurement, and it is here because the same audit lists it beside the others.

Which of these your own pages would fail today

That list is one audit’s wording, not a standard. It is still the most concrete published version of the diagnosing half.

What you can check yourself afterwards

Whatever the report says, the next question is where an improvement would show up, and the answer is narrower than the report implies. One half of an audit produces something you can verify on your own pages; the other produces a number that your own Google reporting will not isolate.

Google documents that its own share of this is not broken out.

Google counts appearances in its AI features inside overall search traffic, under the Web search type. So on that surface there is no line to watch after an audit. What the other engines report to publishers is published separately by each of them and is not covered here. That leaves three checks you can run yourself. The first is repeating the audit’s own prompts yourself a few weeks later and seeing whether the answers moved, which is the procedure in how to measure AI visibility. The second is the diagnostic half: the markup either exists now or it does not, and that is verifiable on your own pages.

Four search types, and the one you are paying to improve is inside the first

Search Console · Performance · Search type filter

1Web

Image

Video

News

  • Google states that appearances in its AI features are counted in overall search traffic, under the Web type. The marked row is where they land.
  • There is no fifth row to select, so no filter on this screen separates an AI appearance from an ordinary blue link.
  • That is why an audit’s promised improvement has no line of its own here. Ask the seller where they expect it to show before you buy.
Our own drawing of the search type filter. The four types are the ones Google offers; the marking is ours.

There is a third check worth naming because it is free and nobody sells it: re-read the report’s own claims about your pages. If it says your product markup is missing, that is verifiable in a browser in a few minutes, and if it says your content does not match how people ask, the prompts it used are either in the report or they are not. An audit whose diagnostic claims cannot be traced to something on your own site has given you an opinion in the shape of a finding. If the report gave you a percentage, the thing to record beside it is the sample it came from. What that percentage is a share of, and why the denominator decides whether it can be compared to anyone else’s, is AI share of voice.

Four ways an audit misleads

The first two come from the same root, which is a one-off reading of a moving measurement presented as a state of affairs. The third comes from the word not naming its contents, and the fourth from expecting a platform to report something it documents that it does not.

A report with no sample size cannot be re-read later.

The only published sample size, and what it adds up to

A published recommendation to start with 20 to 40 prompts total distributed across journey stages, followed by three bullets: 10 to 20 awareness prompts, 20 to 30 consideration prompts, and 5 to 10 brand evaluation prompts tracked separately from category prompts
  1. 1The three lines below the total come to between 35 and 60 prompts. The total above them says 20 to 40, so the headline figure and the split it introduces cannot both hold.
  2. 2Take the number you can act on and keep the arithmetic in view. Whichever of the two you adopt, write it down beside your report, because a figure with no sample cannot be compared to itself next month.
seranking.com, how to choose prompts to track, read 10 September 2026.

The free ones show you what that looks like. A method described in one sentence is the common shape, and the sentence typically settles none of the three: how many prompts, which engines, how many repeats.

The mistakeWhat the evidence saysWhat to do instead
Treating the score in the report as your positionThe same prompts returned the same brand list under one time in a hundredRead it as one reading, and record the date and the sample
Comparing this audit to last quarter’s from another supplierDifferent prompt sets produce different percentages of different totalsCompare only when the set and the counting rule match
Buying a scorecard when you wanted fixesTwo audits have no diagnostic step at allAsk which half the deliverable is before commissioning
Expecting the report to show up in your analyticsGoogle counts AI feature appearances inside the Web search typeVerify the diagnostic half on your own pages instead
The first two are the same mistake at two scales: a reading without its sample cannot be compared, including to itself.

A fifth problem is not yours to solve. Almost every audit here is published by a company selling into this category, and the field measurements come from vendors too. The one thing you can check without trusting anyone is what each audit says it covers, which is what the figure at the top compares.

When an audit is the right purchase

Given that everything above is vendor-published, the safe reading is the narrow one. Nothing here shows that commissioning either half improves anything; what it shows is which half answers which question, so ask for the one whose question you have.

One half asks about your site. The other asks about an engine’s output.

The asymmetry is the whole decision. Whether your product pages carry machine-readable markup is a fact about your site, checkable by you at any time and not dependent on what four engines said last Tuesday. What good markup looks like will keep changing; whether you have any is not a matter of opinion. Whether four engines named you last Tuesday is a draw from a distribution, and a report that does not say how many prompts produced it cannot be checked by you, by them, or by the version of you reading it next quarter.

An audit ends where the work starts. Knowing which half of the question you have does not change either answer.

The markup still has to be written. The pages still have to say something an engine will repeat, and that second job is our answer engine optimization service.

Sources

  1. Despina Gavoyannis, Ahrefs AI visibility audit in eight steps: scope, benchmark, branded and unbranded response analysis, top cited pages, competitor comparison. Ahrefs sells SEO software 30 October 2025, accessed 4 September 2026
  2. Todd Paris, Search Engine Journal The AI search visibility audit: fifteen questions for a CMO, of which seven are published and five diagnose causes on your own pages 28 October 2025, accessed 4 September 2026
  3. Yotpo AI visibility audit in four stages: query mapping, manual prompting, mapping source citations back to their origins, then a stage on automated technical scanning. Yotpo sells ecommerce marketing software 27 July 2026, accessed 4 September 2026
  4. Ariad Partners AI visibility audit described as a diagnostic, delivered within five business days, covering structural and content gaps preventing citations updated 6 August 2026, accessed 4 September 2026
  5. White Label IQ White label AI visibility audit: evidence pack across four engines, scorecard, priority roadmap, and no fixes implemented in this phase. Typically one to two weeks 2026, accessed 4 September 2026
  6. Adamigo Free AI search grader: five checks covering mentions, links, sentiment, competitors and missing topics, probing multiple engines with relevant prompts 26 August 2026, accessed 4 September 2026
  7. SE Ranking How to choose prompts to track: start with 20 to 40 prompts across journey stages, two or three models, and track for at least 30 days before drawing conclusions. SE Ranking sells SEO software 23 April 2026, accessed 4 September 2026
  8. SE Ranking AI search visibility tracker product page, which does not repeat the thirty-day rule from the same company’s guide accessed 4 September 2026 This page carries no publication date of its own.
  9. Amplitude Free AI visibility report: a registration page stating no method, prompt count or engine list accessed 4 September 2026 This page carries no publication date of its own.
  10. Rand Fishkin, SparkToro AIs are highly inconsistent when recommending brands: 12 prompts run a combined 2,961 times, and the same brand list returned under one time in a hundred. SparkToro sells audience research software 28 January 2026, accessed 3 September 2026
  11. Inspeccia Why AI visibility tools disagree: across 167 domains measured twice, the visible ones moved by an average of 30.8 points. Inspeccia sells a visibility product 28 July 2026, accessed 3 September 2026
  12. Profound Is once a day enough: 753 prompts across seven platforms for two weeks, once a day landing within about two points of ten times a day. Profound sells a visibility platform 8 July 2026, accessed 4 September 2026
  13. Greg Jarboe, Search Engine Journal Clovion data: a single follow-up question changed 62 percent of AI brand recommendations 7 July 2026, accessed 3 September 2026
  14. Google Search Central AI features and your website: appearances in AI features are included in overall search traffic within the Web search type last updated 10 December 2025, accessed 4 September 2026
  15. Google Search results for “ai visibility audit” read 4 September 2026

Questions people ask

What is an AI visibility audit?

A one-off report on whether answer engines name your brand, and in some cases on why they do not. What arrives under that name varies. Of the audits published openly, one is entirely measurement and another is mostly diagnosis of your own pages.

The measuring half is the same work as ongoing tracking, run once. The diagnosing half asks whether your pages are structured, verifiable and fresh enough to be quoted. That is a question about your site, not about what an engine answered today.

How long should an AI visibility audit take?

The two answers anyone publishes are far apart. One agency delivers within five business days and another quotes one to two weeks, while a tracking guide from a different vendor says to track for at least thirty days before drawing conclusions.

Both can be right, because they are about different halves. Checking your markup does not need thirty days of observation. Concluding where you appear needs either time or scale: one experiment found a set of 753 prompts read once a day landing within about two points of the same set read ten times, so a large enough one-off may substitute for the month. None of the six audits states a size, so neither route can be confirmed.

Can ChatGPT do an SEO audit?

That is a different question, and it comes up constantly. Using an AI tool to perform audit work and auditing your visibility inside AI answers are two jobs that share a vocabulary, and the split is set out in what AI search optimization is.

Almost nobody writing about audits addresses it.

How many prompts should an audit run?

No audit says. One vendor tracking guide recommends starting at 20 to 40 prompts across journey stages and two or three engines, which is the only published number nearby.

The size matters because a large enough set averages out the movement between runs, and the one published test of that sits at 753 prompts, where reading once a day landed within about two points of reading ten times a day. Nobody has published where the effect starts, so a smaller set is untested, not ruled out.