What an AI visibility audit includes
An audit implies a fixed scope: somebody looks at a defined list and reports what they found. Here the list is not fixed. What you get ranges from a report on where you appear to a diagnosis of why you do not.
Despina Gavoyannis at Ahrefs sets out eight steps and not one of them is diagnostic. Another audit is mostly diagnosis.
An ecommerce platform splits its audit four ways: query mapping, manual prompting, tracing each cited source back to where it came from, then an automated technical scan. That third stage is what tells you why you are missing, not just that you are.
The six audits published openly, read 4 September 2026Which of the two deliverables you are being quoted for
What each audit covers.

One provider is explicit about which half it sells. Its page opens by calling the audit a diagnostic. It lists structural and content gaps preventing citations alongside the presence check, and that is the sentence Google’s own answer quotes when it defines the term. Not every audit sold under that name contains a diagnostic step at all, so the definition Google is repeating fits some of them and misdescribes the rest. The gap is not a quibble about labels, because the two deliverables answer different questions and cost different amounts of work. The measuring half is the same work as how to measure AI visibility, run once instead of continuously. The diagnosing half asks why you are absent. That is a question about your pages, not about the engines. That does not make its findings permanent: what counts as well-structured or as fresh is a moving standard, and one of the published checks asks exactly that.
What one reading can carry
The diagnosing half looks at your own pages, so re-running it does not depend on what an engine answered today. The measuring half is a single reading of something that moves, and that is where the missing sample size in all six starts to matter.
Three separate measurements of how far one reading can move.
There is a condition under which one reading is enough, and it is the strongest defence of the format. Profound ran 753 prompts once a day against ten times a day for two weeks. The two readings landed within about two percentage points, because a set that size averages the movement out. So a one-off audit over a large enough prompt set is a reasonable instrument. Whether any particular audit meets that condition cannot be checked, because none of the six says how many prompts it used.
How long an audit takes: thirty days or five
The clearest way to see the tension in the format is to put two sentences from this category side by side. Both are written by companies selling into it, and they are about the same measurement.
One is an observation window. The other is a delivery promise. They are not the same kind of number, which is exactly why they are worth putting together.
SE Ranking (23 April 2026) and Ariad Partners, read 4 September 2026How long the number under your audit was observed
Two vendor statements about the same kind of measurement.

This does not make the five-day audit dishonest. A diagnosis of your pages asks whether your product information is structured, whether your expertise claims are verifiable, and whether the content matches how people ask. None of that needs thirty days of observation. It is the measuring half that inherits the thirty-day problem, and the audits that are entirely measurement inherit all of it.
Running that diagnosis on your pages, and then doing the work it names, is our AI visibility service.
Four questions to ask before buying one
The useful version of this purchase is specific about which half you are buying and about the sample behind the half that has one. Four questions cover it.
Buy the diagnosis for the fixes, and the measurement for the baseline.
| Ask before you commission | Why it changes what you get | What a non-answer means |
|---|---|---|
| Is this a report or a diagnosis | Two audits contain no diagnostic step, and one is mostly diagnosis | You may receive a scorecard when you wanted a list of fixes |
| How many prompts, and which engines | The count is what decides whether a single reading averages out; the engine list decides which answers were sampled at all | The number cannot be compared to anything, including your own next one |
| Are follow-up questions included | One follow-up changed 62 percent of brand recommendations in one measurement | The audit measured openings, not conversations |
| What happens after the report | One agency states plainly that no fixes are implemented in this phase | The roadmap may be the deliverable, not the work |
The third question is the one nobody asks, and it comes from a measurement. If the prompt set is built entirely from opening questions, the audit has measured how engines answer people who have not asked anything yet. One vendor’s data puts a size on the difference, with a single follow-up question changing 62 percent of the brand recommendations. Whether your own buyers ask a second question is a fact about them, and this measurement does not establish it. What it does establish is that a set of opening questions and a set including follow-ups return substantially different brand lists. On the diagnosing half, the most concrete list published anywhere is Search Engine Journal’s. Whether expertise claims are verifiable. Whether content matches how people query engines. Whether product information carries machine-readable markup, and whether content that looks fresh to you is fresh to an engine. Those are checks on your own pages, not on an engine’s output, so re-running them tomorrow does not depend on what an engine happened to answer. Whether the answer stays useful is a different question: what counts as fresh, or as well-structured, is a judgment that can change.
Search Engine Journal, read 4 September 2026Which of these your own pages would fail today
The questions this audit asks. The last one is measurement, and it is here because the same audit lists it beside the others.

That list is one audit’s wording, not a standard. It is still the most concrete published version of the diagnosing half.
What you can check yourself afterwards
Whatever the report says, the next question is where an improvement would show up, and the answer is narrower than the report implies. One half of an audit produces something you can verify on your own pages; the other produces a number that your own Google reporting will not isolate.
Google documents that its own share of this is not broken out.
Google counts appearances in its AI features inside overall search traffic, under the Web search type. So on that surface there is no line to watch after an audit. What the other engines report to publishers is published separately by each of them and is not covered here. That leaves three checks you can run yourself. The first is repeating the audit’s own prompts yourself a few weeks later and seeing whether the answers moved, which is the procedure in how to measure AI visibility. The second is the diagnostic half: the markup either exists now or it does not, and that is verifiable on your own pages.
Four search types, and the one you are paying to improve is inside the first
Search Console · Performance · Search type filter
1Web
Image
Video
News
- Google states that appearances in its AI features are counted in overall search traffic, under the Web type. The marked row is where they land.
- There is no fifth row to select, so no filter on this screen separates an AI appearance from an ordinary blue link.
- That is why an audit’s promised improvement has no line of its own here. Ask the seller where they expect it to show before you buy.
There is a third check worth naming because it is free and nobody sells it: re-read the report’s own claims about your pages. If it says your product markup is missing, that is verifiable in a browser in a few minutes, and if it says your content does not match how people ask, the prompts it used are either in the report or they are not. An audit whose diagnostic claims cannot be traced to something on your own site has given you an opinion in the shape of a finding. If the report gave you a percentage, the thing to record beside it is the sample it came from. What that percentage is a share of, and why the denominator decides whether it can be compared to anyone else’s, is AI share of voice.
Four ways an audit misleads
The first two come from the same root, which is a one-off reading of a moving measurement presented as a state of affairs. The third comes from the word not naming its contents, and the fourth from expecting a platform to report something it documents that it does not.
A report with no sample size cannot be re-read later.
The only published sample size, and what it adds up to
- 1The three lines below the total come to between 35 and 60 prompts. The total above them says 20 to 40, so the headline figure and the split it introduces cannot both hold.
- 2Take the number you can act on and keep the arithmetic in view. Whichever of the two you adopt, write it down beside your report, because a figure with no sample cannot be compared to itself next month.
The free ones show you what that looks like. A method described in one sentence is the common shape, and the sentence typically settles none of the three: how many prompts, which engines, how many repeats.
| The mistake | What the evidence says | What to do instead |
|---|---|---|
| Treating the score in the report as your position | The same prompts returned the same brand list under one time in a hundred | Read it as one reading, and record the date and the sample |
| Comparing this audit to last quarter’s from another supplier | Different prompt sets produce different percentages of different totals | Compare only when the set and the counting rule match |
| Buying a scorecard when you wanted fixes | Two audits have no diagnostic step at all | Ask which half the deliverable is before commissioning |
| Expecting the report to show up in your analytics | Google counts AI feature appearances inside the Web search type | Verify the diagnostic half on your own pages instead |
A fifth problem is not yours to solve. Almost every audit here is published by a company selling into this category, and the field measurements come from vendors too. The one thing you can check without trusting anyone is what each audit says it covers, which is what the figure at the top compares.
When an audit is the right purchase
Given that everything above is vendor-published, the safe reading is the narrow one. Nothing here shows that commissioning either half improves anything; what it shows is which half answers which question, so ask for the one whose question you have.
One half asks about your site. The other asks about an engine’s output.
The asymmetry is the whole decision. Whether your product pages carry machine-readable markup is a fact about your site, checkable by you at any time and not dependent on what four engines said last Tuesday. What good markup looks like will keep changing; whether you have any is not a matter of opinion. Whether four engines named you last Tuesday is a draw from a distribution, and a report that does not say how many prompts produced it cannot be checked by you, by them, or by the version of you reading it next quarter.
An audit ends where the work starts. Knowing which half of the question you have does not change either answer.
The markup still has to be written. The pages still have to say something an engine will repeat, and that second job is our answer engine optimization service.
Sources
- Despina Gavoyannis, Ahrefs AI visibility audit in eight steps: scope, benchmark, branded and unbranded response analysis, top cited pages, competitor comparison. Ahrefs sells SEO software
- Todd Paris, Search Engine Journal The AI search visibility audit: fifteen questions for a CMO, of which seven are published and five diagnose causes on your own pages
- Yotpo AI visibility audit in four stages: query mapping, manual prompting, mapping source citations back to their origins, then a stage on automated technical scanning. Yotpo sells ecommerce marketing software
- Ariad Partners AI visibility audit described as a diagnostic, delivered within five business days, covering structural and content gaps preventing citations
- White Label IQ White label AI visibility audit: evidence pack across four engines, scorecard, priority roadmap, and no fixes implemented in this phase. Typically one to two weeks
- Adamigo Free AI search grader: five checks covering mentions, links, sentiment, competitors and missing topics, probing multiple engines with relevant prompts
- SE Ranking How to choose prompts to track: start with 20 to 40 prompts across journey stages, two or three models, and track for at least 30 days before drawing conclusions. SE Ranking sells SEO software
- SE Ranking AI search visibility tracker product page, which does not repeat the thirty-day rule from the same company’s guide This page carries no publication date of its own.
- Amplitude Free AI visibility report: a registration page stating no method, prompt count or engine list This page carries no publication date of its own.
- Rand Fishkin, SparkToro AIs are highly inconsistent when recommending brands: 12 prompts run a combined 2,961 times, and the same brand list returned under one time in a hundred. SparkToro sells audience research software
- Inspeccia Why AI visibility tools disagree: across 167 domains measured twice, the visible ones moved by an average of 30.8 points. Inspeccia sells a visibility product
- Profound Is once a day enough: 753 prompts across seven platforms for two weeks, once a day landing within about two points of ten times a day. Profound sells a visibility platform
- Greg Jarboe, Search Engine Journal Clovion data: a single follow-up question changed 62 percent of AI brand recommendations
- Google Search Central AI features and your website: appearances in AI features are included in overall search traffic within the Web search type
- Google Search results for “ai visibility audit”
Questions people ask
What is an AI visibility audit?
A one-off report on whether answer engines name your brand, and in some cases on why they do not. What arrives under that name varies. Of the audits published openly, one is entirely measurement and another is mostly diagnosis of your own pages.
The measuring half is the same work as ongoing tracking, run once. The diagnosing half asks whether your pages are structured, verifiable and fresh enough to be quoted. That is a question about your site, not about what an engine answered today.
How long should an AI visibility audit take?
The two answers anyone publishes are far apart. One agency delivers within five business days and another quotes one to two weeks, while a tracking guide from a different vendor says to track for at least thirty days before drawing conclusions.
Both can be right, because they are about different halves. Checking your markup does not need thirty days of observation. Concluding where you appear needs either time or scale: one experiment found a set of 753 prompts read once a day landing within about two points of the same set read ten times, so a large enough one-off may substitute for the month. None of the six audits states a size, so neither route can be confirmed.
Can ChatGPT do an SEO audit?
That is a different question, and it comes up constantly. Using an AI tool to perform audit work and auditing your visibility inside AI answers are two jobs that share a vocabulary, and the split is set out in what AI search optimization is.
Almost nobody writing about audits addresses it.
How many prompts should an audit run?
No audit says. One vendor tracking guide recommends starting at 20 to 40 prompts across journey stages and two or three engines, which is the only published number nearby.
The size matters because a large enough set averages out the movement between runs, and the one published test of that sits at 753 prompts, where reading once a day landed within about two points of reading ten times a day. Nobody has published where the effect starts, so a smaller set is untested, not ruled out.