Why ChatGPT Picks Your Competitor Before It Reads Your Site
The shortlist exists before your server is ever contacted
You paid for schema markup, an llms.txt file, and a faster server, and ChatGPT still recommends the clinic two streets over. That is not broken work. It is the second half of the job, bought as if it were the first. ChatGPT writes its own search query before it fetches anything, that query already carries brand names, and no change you make to your own website can put a name into it.
The technical AI-visibility work being sold to business owners right now cannot get you onto ChatGPT's shortlist. The shortlist is drawn from the model's memory before your site enters the picture, and only off-site reputation moves that. Search "how does ChatGPT decide which businesses to recommend" and the top pages will hand you a checklist of on-site fixes. Not one of them tells you the order of operations, which is the only detail that decides whether those fixes can help you at all.
The clinic that got named did not out-optimize you on the day of that conversation. It was already in the running before the search ran. Suganthan Mohanadasan tested this and published the results in Search Engine Journal on August 14, 2026, working from 57 conversations and 3,554 labeled pages. In 21 of the 27 test conversations he logged, ChatGPT's first search query already contained brand names the user never typed. The user asked a generic question. The model answered with a list it brought with it.
Query fan-out is the step where ChatGPT rewrites your one question into several searches of its own before it goes looking. Everyone assumed fan-out was how the model discovers who the players are. Mohanadasan's data says otherwise.
"So the fan-out was never a search for candidates; it was ChatGPT going down a list it already had, one name at a time."
Suganthan Mohanadasan, Search Engine Journal, August 14, 2026.
Picture your own business. A physio clinic in Leeds, three locations, decent website, someone in the family runs the Instagram. A woman with a bad shoulder opens ChatGPT and asks it to recommend a physio near her. Before it searches for anything, the model has drafted its own query with two or three clinic names inside it. If yours is not one of them, the rest of that conversation is a formality, and you never see the loss because nothing shows up in your analytics. Silence is not neutral. It is a booked appointment that went somewhere else.
The gap between being named and being merely findable is roughly 33 to 1
Two numbers from the same study carry the whole argument. Brands named inside ChatGPT's own generated query appeared in the final answer 68.9% of the time. Pages that were fetched, read, and never named in that query appeared 2.1% of the time. Mohanadasan's summary: "Being in that query is worth about 33 times more than being findable."
Run the percentages on a made-up hundred. Out of a hundred conversations, a brand inside ChatGPT's own query shows up in about sixty-nine answers. A brand that gets fetched but never named shows up in two. The hundred is mine for illustration; the two percentages are Mohanadasan's. That is the difference between a full diary and a quiet one.
Mohanadasan's dataset is 57 conversations on a single account. That is a real experiment with a small sample, not a census of the internet. It points in the same direction as the wider data below, and I would act on it, but I would not quote it as a law of physics.
You were sold Game Two and told it was Game One
Split the work into two games and the whole confused market gets clear. Game One is getting onto the shortlist the model brings to the conversation. Game Two is getting quoted, cited, and described accurately once you are already on it. Almost everything the AI-visibility industry currently sells is Game Two.
Schema markup is code in your page that labels what things are, so a machine knows this string is a price and that one is an address. An llms.txt file is a plain text file you host that tells AI crawlers what your site covers. Both are reasonable. Neither can be read by anything until a request reaches your server, and by then the list is written. Mohanadasan puts the structural problem in one line.
Olivier de Segonzac's breakdown of ChatGPT's retrieval stack in Search Engine Land on August 17, 2026 adds the mechanical reason this holds even later in the process. Retrieval runs through an index and a cache, which are stored copies of pages the system keeps on hand instead of visiting your site fresh each time, so the model often reads a stored copy of your page, not a fresh visit to your web host. The live tuning you did last Tuesday may not be the version being read.
I have watched this pattern since 2008, and more than once I have watched an old lever get resold as a new one. The tell is always the same. A vendor takes work that has a genuine but narrow effect, attaches it to the outcome you actually want, and prices it against that outcome. Ask what specifically changes when your llms.txt file goes live, and the answer will be about compliance with a format, never about a customer hearing your name.
| The work | Which game | What it can and cannot change |
|---|---|---|
| Schema markup | Game Two | Helps a machine read your facts correctly once your page is fetched. Cannot put your name in the query that ran before the fetch. |
| llms.txt file | Game Two | Requires your server to be contacted. That contact happens after the shortlist is drawn. |
| Page speed and Core Web Vitals | Game Two | Real value for human visitors and classic search, worth doing on its own merits, not as a route onto the list. |
| Clear headings and answer-shaped content | Game Two | Improves your odds inside the 3.1% of retrieved pages that get cited. Only pays out if you were retrieved at all. |
| Being written about, reviewed and compared off-site | Game One | The only category of work that operates on the association the model already carries. Slow, cumulative, cannot be bought as a package. |
| Named customer stories and public reviews | Game One | Puts your name next to your category in text other people wrote, which is what the recall is built from. |
The wider data agrees: recall is built off your site, not on it
One small study would not be enough to change how you spend money. It does not stand alone. Kelsey Libert of Fractl published an AI visibility index in Search Engine Land on August 17, 2026, built from 4,320 responses covering more than 8,500 brands across GPT-4o, Gemini 2.5 and Claude Sonnet 4.6. Only 11% of brands were referenced by all three models. 77% appeared in just one.
Libert's conclusion lands on the same point Mohanadasan reached from the retrieval side.
"What a brand says about itself matters less for AI recall than what the rest of the web has repeatedly said about it."
Kelsey Libert, Fractl, Search Engine Land, August 17, 2026.
Ahrefs studied 75,000 brands in 2025 and found brand mentions correlate 0.664 with AI visibility, with the top quarter of brands by mention volume picking up roughly ten times more AI Overview citations. Correlation is not proof of cause. As a budgeting signal it is still the clearest one available: the thing that moves with AI visibility is how often other people write your name, not how tidy your markup is.
It is worth knowing where the academic work sits in all this. The Princeton and Georgia Tech generative engine optimization study presented at KDD in 2024 examines how to make a source more likely to be used once it is already in the retrieved set. It is useful research and it is entirely Game Two, which does not stop vendors citing it as if it described how you get chosen in the first place.
Getting onto the list is public work you mostly do off your own website
The fix runs on quarters, not a sprint.
"What seems to build it is slow and unglamorous. Being written about, reviewed, compared, and argued over across the open web for years, until the association exists in the training data."
Suganthan Mohanadasan, Search Engine Journal, August 14, 2026. Training data is the enormous body of public text a model learned from before it ever met you.
That is the part I can speak to from my own side of the desk. Across 300-plus businesses in the USA, Canada, the UK, Singapore, Australia and New Zealand, in medical, ecommerce, local services, fintech, education and real estate, the ones who became the obvious name in their category got there by being present in other people's content for years. The London ADHD clinic I booked solid for three straight months, to the point they hired more specialists and outsourced the overflow, did not win on markup. Demand did that. Demand is built in public.
I would run five moves, in this order.
If you can only do one thing, do the third. Being present in the places where your category gets discussed by other people is why your own blog is rarely where AI actually finds you, and it is the fastest available route from invisible to occasionally named. Once names start appearing, the next decision is which single category you want to own, because spreading the effort across five is how small businesses stay off every list. I wrote about why AI keeps naming the same brands and how narrow your window is for exactly that choice.
Measure recall across repeat runs, and ignore the crawler-hit dashboards
The vanity metric in this category is bot traffic. Someone shows you a chart of AI crawler visits to your site and calls it AI visibility. Those hits happen after retrieval, which is after the list was drawn, and with 3.1% of retrieved pages getting cited in Mohanadasan's data, a fetch on its own tells you almost nothing. Schema validation passes and the existence of an llms.txt file are worse: they measure whether a file exists, not whether a customer heard your name.
Three things are worth tracking. First, how often your business appears across repeated runs of the same buyer question, not once, because asking ChatGPT whether it recommends you a single time tells you nothing. Second, the same question across at least three assistants, since Fractl found 77% of brands surfaced in only one model. Third, your count of mentions across the open web, quarter on quarter, which is the closest available proxy for the thing that actually predicts visibility. Search your business name in quotes on Google, set the date filter to the last three months, and write the result count in a spreadsheet on the same day each quarter. The direction matters, not the exact figure.
Set the review to quarterly and hold your nerve between reviews. Off-site recall does not move in a fortnight, and checking it weekly turns a slow compounding asset into a source of anxiety. The number I would expect to move first is mentions. Appearances in answers follow, later, unevenly, and never on the schedule a vendor promised you.
That schedule is also how you spot the sales pitch. Anything advertised as a thirty-day route into ChatGPT's recommendations is Game Two work wearing Game One's clothes, because the mechanism that would make it true does not exist. Nobody can insert your name into a model's memory on a payment plan.
Frequently Asked Questions
Can I pay to get my business into ChatGPT's recommendations?
No, and any service selling that is selling something else. The shortlist ChatGPT carries into a conversation is built from what the open web has said about brands over a long period, so there is no account, fee or file that inserts you into it. The emails offering ChatGPT inclusion are link-buying in a new outfit, and paid link schemes carry their own risks with classic search. Spend the same money on being genuinely quoted, reviewed and compared in places your buyers already read.
How long before ChatGPT starts naming my business?
Quarters, not weeks, and the timeline depends on three things you can check today. How crowded your category is, how much public text already exists with your name in it, and how consistently you can earn new mentions each month. A local service in a thin category shows up faster than a business competing against national brands with a decade of press. Track mentions monthly and answer appearances quarterly, because the mentions move first and the appearances follow.
So was the schema and llms.txt work a waste of money?
Not a waste, just mispriced and misdescribed. That work helps a model read and quote you correctly once your page has been retrieved, which matters given only 3.1% of retrieved pages got cited in Suganthan Mohanadasan's study. It simply cannot put you on the shortlist, because your server is not contacted until after the shortlist exists. Keep it, stop paying premium rates for it, and move the difference into being written about off your own site.
Most owners reading this will do nothing, because the honest version of the work has no dashboard and no thirty-day win. The ones who start now will be the names inside ChatGPT's query in two years, and the rest will keep optimizing pages that never get read. If you want a second opinion on which half of the job your current spend is buying, book a call with me at https://cal.com/johntalaguit/ai-visibility-call and bring your last invoice. Look at it before the call and ask yourself which line item on it has ever caused another human being to write your business name in public.