Why LLM SEO has no positions to rank in
The word SEO carries an expectation that there is a list and a position in it. Say you want to be named when somebody asks Claude about your category. No list, no position. Tactics written for a ranked list transfer badly here.
What happens instead is retrieval. The model issues a search, gets documents back, and writes an answer from them. Whether your page can be in that set is a separate question from how good the page is, and it is answered before any content decision matters.
That shows up in what a change can do for you. On a ranked list, better work moves you from eighth to fourth and you can watch it happen. In an answer, the page is either among the documents the model retrieved or it is not, and if it is, the model decides how much of it to use. No position exists to improve, so a tactic borrowed from ranking work often has nothing to attach to.
How the term arrives, and the verb it brings
- 1Ranking in AI search is in the subtitle. The verb is the problem: it imports an expectation from a ranked list onto a surface that has no positions.
- 2The opening line lists four names for the same work, which is a fair description of the category and also the reason a buyer cannot tell what is being sold.
- 3Marketer Milk is thorough on content tactics. What it does not open is the file that decides whether any of them get read.
In plainer terms, and about its own surfaces, Google puts it this way: its AI features are rooted in the core Search ranking systems and retrieve from the Search index. That is why AI SEO treats being found by an assistant as ordinary search work with one extra file to check.
One brand-level count puts a number on it. Semrush looked at 131 Petlibro pages cited in ChatGPT answers and found 85 percent of them also rank for at least one Google keyword, averaging 19 keywords each.
That overlap is about the Google index, not about the crawler rules. It says cited pages tend to be pages search already knows; it does not say a robots.txt line is what admitted them.
The content half of your work has its own name and its own test. Generative engine optimization was named by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande at KDD 2024.
C-SEO Bench, by Puerto, Gubri, Green, Oh and Yun, later ran those edits wider and found three of fifty-four combinations improved. Your crawler settings decide whether any of it reaches your reader at all.
A win on one provider does not carry to the next. Orbit Media Studios read 13,184 citations across 1,765 answers and found all four models naming the same domain for your question in 30 of 1,792 combinations.
Say your category question is answered every day by an assistant and your brand is never in it. The first thing to establish is not whether your pages read well. Ask instead whether the provider can fetch them at all. That answer lives in a file you already own.
Three providers and three crawlers
Each provider runs more than one crawler and gives them different jobs. The distinction that matters is between the crawler that collects training data and the crawler that fetches pages for search answers, because blocking the second removes you from the answers.
OpenAI names OAI-SearchBot as the crawler for search. Its documentation states that sites opted out will not be shown in ChatGPT search answers, though they can still appear as navigational links. A robots.txt change takes about twenty-four hours to register.
The setting that decides ChatGPT eligibility
- 1Each setting is independent of the others. A site can allow the search crawler while disallowing the training one, and OpenAI says so on the page.
- 2The consequence is stated, not implied: opted out means not shown in ChatGPT search answers.
- 3Twenty-four hours from a robots.txt change. That is the feedback loop on this setting, which is worth knowing before you go looking for an effect.
Anthropic publishes a table with three bots and, for each, what happens when you disable it. Claude-SearchBot is the one tied to search answers. Anthropic states that disabling it prevents indexing for search optimisation and may reduce a site’s visibility in user-directed web search.
Three bots, and the consequence spelled out for each
- 1The third column is unusual and it is the useful one. Each row says what you lose by blocking that bot, so the trade is written down, not inferred.
- 2ClaudeBot is training. Claude-User is a fetch made because a person asked. Claude-SearchBot is the one that feeds search answers.
- 3Blocking the first two changes nothing about search visibility. Blocking the third is the decision that does.
The three crawlers, side by side
| Provider | Search crawler | Training crawler | Published consequence of blocking search |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | Not shown in ChatGPT search answers, though still possible as a navigational link |
| Anthropic | Claude-SearchBot | ClaudeBot | May reduce visibility and accuracy in user-directed web search |
| Perplexity | PerplexityBot | Not named on this page | Not stated on the page; the crawler is documented as the one that indexes |
Perplexity publishes its crawlers the same way, separates the scheduled crawler from the one that fetches a page because a user asked, and publishes IP ranges for verification. Its documentation also notes that a web application firewall can block these agents without anybody intending it. Your robots.txt will not show you that.
A second place the block can happen
- 1Two agents with different jobs, and the page says user-initiated requests are not used for indexing.
- 2Published IP ranges mean a claimed crawler can be verified, not trusted, which matters when logs are the evidence.
- 3The firewall section is the part most sites miss. A rule nobody wrote for this purpose can remove you from a surface, and your robots.txt will look fine.
How to check all three in twenty minutes
This is a settings audit, not a content project. The audit is cheap, it changes more than anything else on this page, and most sites have never done it because no guide told them to.
-
Open yourstore.com/robots.txt in a browser
Read it yourself, not through a tool. The file is short and the whole decision is visible in it.
-
Search it for the three search crawlers
OAI-SearchBot, Claude-SearchBot and PerplexityBot. What a Disallow costs differs by provider. OpenAI states outright that opted-out sites will not appear in ChatGPT search answers. Anthropic says it may reduce visibility, and Perplexity does not say.
-
Separate them from the training crawlers
GPTBot, ClaudeBot and the user-initiated agents are different settings with different consequences. Blocking training while allowing search is a coherent position and the platforms support it.
-
Check your firewall and bot rules, not just robots.txt
Perplexity documents this case directly. A rule written to stop scrapers can stop these agents, and nothing in your robots.txt will reveal it. Server logs will.
-
Wait, then look again
OpenAI publishes about twenty-four hours for a robots.txt change to register. Do not judge a change before its own stated window has passed.
The same two words, and three of these six lines decide whether you appear
yourstore.com/robots.txt
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
1User-agent: OAI-SearchBot
Disallow: /
2User-agent: Claude-SearchBot
Disallow: /
3User-agent: PerplexityBot
Disallow: /
- The three marked names are the search crawlers. OpenAI states that an opted-out site will not appear in ChatGPT search answers, Anthropic says visibility may drop, and Perplexity does not say.
- The greyed names are training crawlers. Blocking those and allowing the three above is a coherent position, and all three providers support it.
- Every line here reads Disallow. Which of them costs you an answer depends entirely on which name sits above it.
Only after that does content work make sense, because until the retrieval step can reach the page, nothing written on it is reachable. What that content work is worth is a separate and less settled question, and generative engine optimization sets out what has been tested.
Where the money goes
Nothing in the section above costs money. The spending starts when somebody proposes to change the content, and that is where the price ranges from nothing to a full retainer.
Editing robots.txt is a developer task. Checking firewall rules is the same. Neither has a subscription attached and neither is sold by anybody. That is part of why they are missing from guides written to sell something.
That absence should be named, not glossed. The single most decisive action in this whole category is free, takes twenty minutes, and has no vendor attached to it. Advice written to sell something will tend to skip it, and the skipping is not dishonesty so much as the ordinary shape of a market.
Measurement is where the tools appear. Trackers that ask assistants your questions start at $29 a month with Otterly.AI and run past $250 with Scrunch AI, which also watches agent traffic hitting your site. AI visibility tools compares three of them on coverage and sampling.
The content work itself is a normal search budget with a different reason attached. Google’s guidance for its own AI features asks for the same things ordinary search work asks for, and says so explicitly.
Say you are quoted for a programme built around this term. Ask which part of it is the settings audit, and what the rest is buying on top. A quote that cannot separate the two is pricing ordinary search work under a newer name. That may still be a good buy, but compare it against ordinary search prices, not against a category with no benchmark.
Say you want the second half handled and not only priced: that is our LLM optimization service.
Four LLM SEO mistakes the term produces
Most of that bill goes on mistakes nobody had to make. Each of these follows from treating the work as ranking. They cost different amounts and the first one is the only one that can remove you from a surface completely.
Blocking the search crawler while meaning to block the training crawler. The names are similar, the settings are independent, and the consequence is not symmetrical. One protects your content from a training set. The other removes you from the answers.
Adding a file the platforms do not read. Google states plainly that you do not need to create new machine-readable files, AI text files or markup. Search does not use them.
The list Google publishes of things to ignore
- 1Five items, and each one is sold as a service somewhere. The heading names the category the advice belongs to.
- 2The llms.txt line is the bluntest: you do not need to create new machine-readable files, and Google Search itself does not use them.
- 3The structured data line is more careful and worth reading exactly: not required for generative AI search, no special markup, and still a good idea for rich results.
Chunking pages into fragments so a model can read them. Google says there is no requirement to break content into tiny pieces and no ideal page length. The instruction produces worse pages for people and buys nothing measurable.
Judging a change before its window has passed. OpenAI publishes twenty-four hours for robots.txt. Answers also vary between runs, so a single check after a change is an anecdote, not a result.
Which two questions your server can answer
So check more than once. Your own server answers the first two questions. The third is the one everybody wants, and no report anywhere carries it.
Your logs answer the first question. If OAI-SearchBot, Claude-SearchBot or PerplexityBot has fetched a page, the request is in them, with a user agent and an IP you can verify against the ranges each provider publishes.
Do that check before anything else, because it is evidence and everything else here is inference. If a provider’s search crawler has never appeared in your logs, no content decision explains that, and a tracker will not either.
Rule out the ordinary causes before you conclude you are blocked. Short log retention, a firewall dropping the request before it is written, and a failed IP check all look identical to never being fetched.
If it appears regularly, the eligibility question is answered. What gets chosen after that is the next question, and no provider has published it.
Search Console answers the second, partly. Google reports AI feature traffic inside overall Search under the Web search type. A click from an AI Overview is in your total; an appearance that produced no click is not.
Mangools publishes a free AI Search Grader covering three models, scoring visibility as the share of prompts where a brand appears in the top twenty. Start there before a subscription.
The third question, whether an answer named you, has no report. You answer it by asking, either by hand or by paying something to ask on a schedule, and how to measure AI visibility sets out the free version.
Treat those three rows as three separate reports. A brand can be fetched daily, receive no clicks that any tool can attribute, and still be named in answers people act on. Reconciling them into one funnel produces a number wrong in a direction you cannot check.
Where each question gets answered
- 1 Did the crawler fetch the page Server logs. Verifiable against the IP ranges each provider publishes.
- 2 Did a click arrive from an AI feature Search Console, folded into Web search totals and not separable.
- 3 Did an answer name the brand Only by asking. No report exists, on any surface.
Read together, the three rows return you to where this page started. The one lever with a documented consequence is a crawler line in a single file, and everything sold above it is a report about what happens afterwards.
That is the whole shape of the category today. Settle the file first. Measure the two questions your own server answers, and treat the third as a question you ask, not a number you buy.
Sources
- OpenAI Overview of OpenAI Crawlers: OAI-SearchBot surfaces websites in ChatGPT search results; sites opted out will not be shown in ChatGPT search answers; robots.txt changes take about 24 hours to register
- Anthropic Does Anthropic crawl data from the web: publishes ClaudeBot, Claude-User and Claude-SearchBot with a column stating what happens when each is disabled
- Perplexity Perplexity Crawlers: separates PerplexityBot from user-initiated fetches, publishes IP ranges, and documents web application firewall configuration
- Google Search Central Optimizing your website for generative AI features: mythbusting section names llms.txt and special markup, chunking, rewriting for AI, inauthentic mentions and overfocusing on structured data as things to ignore
- Google Search Central Same guide: generative AI features are rooted in the core Search ranking and quality systems, using retrieval-augmented generation and query fan-out
- Google Search Central AI features and your website: no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary
- Google Search Central Search Console documentation: AI feature traffic is reported inside overall Search traffic under the Web search type
- Omar Uddin, Marketer Milk LLM SEO: The ultimate guide to ranking in AI search, October 2025
- Carlos Silva, Semrush ChatGPT SEO: of 131 Petlibro pages cited in ChatGPT answers, 85 percent also rank for at least one keyword in Google, with an average of 19 keywords per cited page
- Puerto, Gubri, Green, Oh and Yun C-SEO Bench, NeurIPS Datasets and Benchmarks 2025: three of fifty-four task and domain combinations improved on making no edit
- Orbit Media Studios 13,184 citations across 1,765 answers: all four models cited the same domain for the same question in 30 of 1,792 query and domain combinations
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande GEO: Generative Engine Optimization, KDD 2024: the paper that named the method and measured share of a generated answer
- Mangools AI Search Grader: a free tool covering three AI models without payment, defining visibility as the share of prompts where a brand appears in the top 20
- Otterly.AI Pricing: entry plan at $29 a month with 15 search prompts
- Scrunch AI Pricing: Core plan at $250 a month with 125 unique prompts and five site audits
Questions people ask
What is the difference between LLM SEO and GEO SEO?
In practice, the name. Both describe making a site more likely to be named inside AI answers, and the products sold under each label are largely the same.
One difference matters: generative engine optimization is the name of a specific method from a 2024 academic paper, so it points at something testable. LLM SEO is an industry coinage and points at nothing in particular.
Does llms.txt help with LLM SEO?
Google states directly that you do not need to create new machine-readable files or AI text files to appear in its features, and that Search does not use them.
No other provider names the file in its crawler documentation either. Publishing one costs almost nothing and no published evidence shows it changes anything. Treat it as optional, not as a step.
Should I block AI crawlers?
That depends which one, and the providers make the distinction for you. Blocking a training crawler keeps your content out of a training set and leaves search visibility alone.
Blocking a search crawler is different. OpenAI says opted-out sites will not be shown in ChatGPT search answers, and Anthropic says disabling Claude-SearchBot may reduce visibility in user-directed web search. Those are the settings to be deliberate about.
How long does a robots.txt change take to work?
OpenAI publishes about twenty-four hours from a robots.txt update for its systems to adjust. The other providers do not publish a figure.
Wait past the stated window before judging, and remember that answers vary between runs, so one check after a change tells you very little on its own.