AI Search Has Five Stages. Your Strategy Covers One.
If your content strategy is still "publish good articles and hope Google ranks them," you are optimizing for one stage of a five-stage machine. Before anything you write can appear in a ChatGPT answer, a Google AI Overview, or a Perplexity result, it has to survive crawling, indexing, retrieval, synthesis, and presentation. Fail at any one of those stages and you are invisible, no matter how good the writing is. Most SEO strategies, including the one your agency is probably running, only address the first stage, and that is why so many businesses with decent content never show up in AI answers.

The Five Stages Between Your Content and an AI Answer
Every AI search system works roughly the same way under the hood, whether it is Google's AI Overviews, ChatGPT with search, or Perplexity. Your content passes through a pipeline, and each stage filters out most of what enters it. Here is each stage in plain English, along with what failure looks like.
Stage 1: Crawling and ingestion
Before an AI system can use your content, a bot has to fetch it. OpenAI sends GPTBot, Anthropic sends ClaudeBot, Perplexity sends PerplexityBot, and Google uses its existing crawl infrastructure. This is the stage everyone already understands, because it is the same idea as classic SEO crawling. But there is a catch most owners have never heard of. Vercel analyzed AI crawler behavior, including more than 500 million GPTBot fetches, and found no evidence of JavaScript execution. These bots grab the raw HTML and leave. If your site is built so that the actual text only appears after JavaScript runs in a browser, which is common with modern site builders and web apps, the AI crawlers see an empty shell. Google can render JavaScript. Most AI bots cannot.
What this means for your business: your site can look perfect to you and be functionally blank to the systems your buyers are asking for recommendations.
Stage 2: Indexing and embedding
Once fetched, your content gets stored, and for AI systems it gets converted into embeddings, which are mathematical representations of meaning. Your pages also get split into chunks, passage-sized pieces that the system handles independently. This is a real change from keyword-era indexing. The machine is no longer matching the words you used. It is matching what your passage means, one chunk at a time, often without the rest of the page attached.
What this means for your business: a page where every section makes sense on its own gets represented accurately. A page where the point only emerges after reading 1,500 words of wind-up gets chopped into chunks that individually say nothing.
Stage 3: Retrieval
When someone asks a question, the system does not read the whole internet. It pulls a shortlist of relevant chunks from its index, and only those make it to the model. Google's AI Mode does this through a technique called query fan-out, which Google has discussed publicly and which Ahrefs and others have documented: the system breaks one question into many sub-queries and runs them in parallel. Ask "best accounting software for a small construction firm" and the system may quietly also search pricing, integrations, reviews, alternatives, and industry fit, then merge the results. Your content competes in each of those hidden searches separately.
What this means for your business: if you only have content for the main question and nothing covering the surrounding sub-questions, you lose most of the hidden contests and may never enter the answer at all.
Stage 4: Synthesis and citation selection
The model now reads the retrieved chunks and writes an answer, deciding which sources to lean on and which to cite. This stage rewards content that is clear, specific, and corroborated by other sources, and it punishes vague claims that nothing else on the web backs up. It is also where platforms diverge sharply. Profound analyzed 680 million citations across ChatGPT, Google AI Overviews, and Perplexity between August 2024 and June 2025 and found very different preferences: Wikipedia is ChatGPT's most-cited domain, Reddit leads on Perplexity, and Google AI Overviews spread citations across a more balanced mix, with Reddit narrowly ahead of YouTube at the top. Strikingly, only around 11 percent of cited domains overlapped between ChatGPT and Perplexity.
What this means for your business: being visible on one AI platform tells you almost nothing about the others. They are different judges with different taste.
Stage 5: Presentation
Finally, the answer reaches a human, and the only question that matters is how you appear in it. Are you named as a recommendation? Cited as a footnote link? Paraphrased without credit? Absent? Semrush's analysis of more than 10 million keywords found AI Overviews stabilizing at around 16 percent of Google queries by late 2025, after peaking near 25 percent in July. And Pew Research Center's behavioral study found that when an AI summary appears, only 8 percent of users click any traditional result, and just 1 percent click the sources cited inside the summary. The mention is the prize, and a mention is not yet the same thing as being believed. The click is a bonus you mostly will not get.
What this means for your business: at this stage the goal is to be the brand the answer names, because almost nobody is clicking through to find you on their own.
Why Most Strategies Only Address One Stage
The SEO industry built its entire playbook for a world with one pipeline: Google crawls, Google indexes, Google ranks, humans click. Every standard deliverable maps to that world. Keyword research, on-page optimization, link building, rank tracking. All of it aims at getting crawled and ranked, which is roughly stage one and a bit of stage two in the new pipeline.
There is also a reporting problem driving this. Agencies sell what they can chart, and rankings are easy to chart. There is no simple dashboard for "your content gets chunked badly" or "you lost the query fan-out on pricing questions." So those failures go unreported, and unreported problems never get budget. The result is a strange situation I see constantly: a business with solid rankings and falling relevance, because the scoreboard they watch only covers the first stage of the game.
The mechanism of failure is worth spelling out, because it explains why "we publish great content" is not a strategy. Your content can be crawled but not indexed usefully, because it is one giant wall of text mixing six topics. It can be indexed but never retrieved, because it answers the headline question and ignores the sub-questions the fan-out actually runs. It can be retrieved but never cited, because the model finds nothing on the wider web corroborating your claims, so it leans on a competitor that reviewers, journalists, and Reddit threads keep mentioning. And it can be cited but not named, a footnote link almost nobody clicks. Each of those is a different problem with a different fix. Treating them all as "we need more content" is like treating every engine noise with more fuel.
What the Data Actually Shows
This is not theory. Each stage has measurable evidence behind it, and the numbers are blunt.
On crawling: alongside Vercel's finding that GPTBot shows zero evidence of JavaScript execution across 500 million fetches, Cloudflare's Radar data shows the sheer asymmetry of AI crawling. In mid-2025, Cloudflare measured Anthropic crawling tens of thousands of pages for every one human visitor it referred, with ratios around 38,000 to 1 in July 2025, while OpenAI sat near 1,100 to 1. AI systems are reading the web at industrial scale and sending back comparatively few clicks. Your content is being consumed by machines whether you planned for it or not, and the payback comes as mentions and influence, not traffic.
On retrieval and synthesis: a separate Profound study of 100,000 prompts shows the platforms barely agree on sources, with only 11 percent domain overlap between ChatGPT and Perplexity. Semrush's separate analysis of Reddit found it among the top-cited domains across multiple AI engines, which tells you community discussion now functions as source material, not just chatter.
On what earns citations: Ahrefs studied 75,000 brands and found brand web mentions correlate with AI Overview brand visibility at 0.664, while backlinks correlate at just 0.218. Seer Interactive's research found brands ranking on page one of Google showed a correlation around 0.65 with being mentioned by large language models, while backlinks showed weak to neutral impact. Two independent teams, same conclusion: the link-building metric the industry spent twenty years optimizing matters far less to AI visibility than how often and how consistently your brand is mentioned across the web.
On why any of this is worth the effort: Semrush's traffic research found the average visitor arriving from AI search converts at a rate that makes them about 4.4 times as valuable as the average organic search visitor, and ChatGPT referral traffic grew 206 percent across 2025. The volume is smaller, the intent is dramatically higher. These are people who already got the full comparison from the AI and are coming to act.
Put together, the data describes a pipeline where stage one is table stakes, stages two and three are won with structure and coverage, and stages four and five are won mostly off your own website, through corroboration and mentions. A strategy that stops at "publish and build links" addresses almost none of that.
How to Fix It, Stage by Stage
You do not need to understand embeddings to act on this. Here are five moves, one per stage, each of which you can do yourself or hand to your web person with a clear instruction.
1. Make sure the bots can actually read your site. The test takes two minutes: open any important page, right-click, choose "view page source," and search for a sentence from your main content. If it is not there in the raw HTML, AI crawlers cannot see it, and you need your developer to enable server-side rendering or switch how the page is built. While you are at it, have them check robots.txt to confirm you are not accidentally blocking GPTBot, ClaudeBot, or PerplexityBot. Blocking them is a legitimate choice for some publishers, but it should be a decision, not an accident.
2. Structure every page so each section stands alone. Write for the chunk, not just the page. One idea per section. Descriptive subheadings that say what the section actually answers, not clever wordplay. Lead each section with the answer, then explain. Add a genuine FAQ where it fits, and use schema markup so machines get explicit signals about what the page contains. This is the cheapest fix on this list and most sites have not done it.
3. Cover the fan-out, not just the keyword. Take your most commercially important topic and list every question a buyer would ask around it: cost, comparisons, alternatives, "is it worth it," industry-specific fit, implementation, risks. That list approximates the sub-queries AI systems generate. Build content that answers each one specifically and honestly, including the comparison pages most businesses are too nervous to publish. You are not writing for one ranking anymore. You are entering a dozen parallel retrieval contests.
4. Earn corroboration off your own site. The citation-selection stage trusts what multiple independent sources agree on. That means industry publications, directories, review platforms, podcasts, local press, and yes, Reddit and community forums where your category gets discussed. Make sure the basic facts about your business are identical everywhere, because conflicting information makes you harder to confidently cite. Given the Ahrefs numbers, a mention in a trade publication is now plausibly worth more to your AI visibility than another backlink from a generic blog.
5. Audit your presentation monthly, per platform. Build a list of ten questions your buyers actually ask, run them through ChatGPT, Perplexity, and Google with AI Overviews, and log the results: named, cited, paraphrased, or absent, and which competitors appear. Because platforms barely overlap in their sources, check each one separately. Thirty minutes a month gives you a visibility scorecard most of your competitors do not have, and it tells you which stage to work on next.
What to Measure and When to Expect Results
Measure the pipeline, not just the old scoreboard. At the crawling stage, have whoever manages your site check server logs or Cloudflare analytics for visits from GPTBot, ClaudeBot, and PerplexityBot; if they are not fetching your pages, nothing downstream can happen. At the presentation stage, track your mention rate from the monthly prompt audit: out of your ten standard questions, how many answers name you? That single number is the closest thing to a rank tracker for AI search. Alongside it, watch AI referral traffic in your analytics (visits from chatgpt.com, perplexity.ai, and similar), branded search volume in Search Console, and the close rate of leads who arrive through those channels, because the Semrush data says they should convert noticeably better.
Timelines vary by stage. Technical fixes, rendering and robots.txt, take effect within weeks as crawlers revisit. Restructured pages can start appearing in retrieved answers within one to three months. Off-site corroboration is the slow lever: expect three to six months before mentions accumulate enough to shift how often AI systems name you, and six to twelve months for the compounding effect to be obvious. If someone promises top-of-ChatGPT placement in thirty days, walk away.
And know the vanity traps. Total traffic is the big one, since AI answers satisfy more queries without a click; your traffic can stay flat while your influence grows. Keyword rankings alone are the second, since ranking number one means little on the 16 percent of queries where an AI Overview answers first. The third trap is obsessing over whether one specific prompt mentions you today, because AI answers vary run to run. Track your mention rate across a consistent basket of prompts over months, the way you would track an average, not a single coin flip.
These five stages live inside a larger system, and I lay the whole thing out in my complete guide to AI search visibility.
Frequently Asked Questions
Do I need separate content for AI search and regular Google SEO?
Mostly no, and be suspicious of anyone selling you a separate "AI content" package. The same qualities win both: clear structure, sections that stand alone, specific verifiable claims, real expertise, and coverage of the surrounding questions buyers ask. Seer Interactive's research found page-one Google rankings correlate strongly with being mentioned by AI tools, around 0.65, so classic SEO and AI visibility largely reinforce each other. The genuine additions are technical, like server-side rendering and allowing AI crawlers, plus a heavier emphasis on off-site mentions and corroboration.
Should I block AI crawlers from my website?
For most businesses, no. Cloudflare's data shows the trade is real: AI platforms crawl thousands of pages for every visitor they send, so if you are a publisher whose product is the content itself, blocking can be a rational negotiating position. But if your website exists to win customers, blocking GPTBot, ClaudeBot, and PerplexityBot removes you from the research tools that 6sense found 94 percent of B2B buyers now use. You would be hiding from your own pipeline. Decide deliberately, and for most owners the answer is to stay visible.
How long does it take to show up in AI answers?
It depends entirely on which stage is failing, which is why the audit comes first. If your content is invisible to crawlers because of JavaScript rendering, a technical fix can change things in weeks. If your pages are readable but poorly structured, restructuring typically shows results in one to three months as systems re-fetch and re-process. If your real gap is corroboration, meaning few independent sites mention you, expect three to six months of consistent PR, reviews, and community presence before AI systems start naming you with any regularity. There is no honest universal timeline, only a stage-by-stage diagnosis.
The five stages are not a prediction about where search is going. They are a description of how AI search already works, backed by crawl logs, citation studies, and behavioral data. Your content is being filtered through this pipeline today, and every stage you ignore is a place you silently lose. Diagnose where you are failing, fix that stage first, and measure mentions like you used to measure rankings.