Skip to content
Crawlable

What is Generative Engine Optimization (GEO)? The Complete 2026 Guide

How ChatGPT Search, Perplexity and Claude find, read and cite pages, why that differs from Googlebot, and what the research says actually improves AI visibility.

By Crawlable · · 8 min read

On this page

Generative Engine Optimization (GEO) is the practice of making your pages easy for AI answer engines to find, read, and cite. Those engines include ChatGPT Search, Perplexity, Claude, and Google's AI Overviews. Classic SEO works to win one of ten blue links. GEO works to become one of the handful of sources an AI model quotes when it writes a single answer.

The term is younger than most of the advice written about it. This guide covers where it came from, how answer engines actually work, how they differ from Googlebot, and what the evidence says helps. It also marks clearly where the evidence runs out.

Where the term comes from

"Generative Engine Optimization" was coined in a paper by Pranjal Aggarwal and colleagues, first posted in late 2023 and published at KDD 2024. The authors define a generative engine as a system that retrieves relevant documents and uses a large language model to "generate a response grounded on the sources." They then asked a question that SEO had never needed to ask: if the answer is a paragraph rather than a ranked list, what does "visibility" even mean?

Their answer was two new metrics. The first is position-adjusted word count: how much of the generated answer is drawn from your page, weighted by how early it appears. The second is subjective impression: a model-judged score for how relevant, influential and unique your contribution to the answer was. They built a 10,000-query benchmark (GEO-bench), rewrote source pages in different ways, and measured what changed.

You'll also see the same idea called AEO (answer engine optimization), LLM SEO and AI SEO. They describe the same goal.

How an answer engine builds an answer

Every major answer engine follows roughly the same pipeline. Each stage is a place where a page can drop out.

  1. The question is rewritten. OpenAI says ChatGPT search "typically rewrites your query into one or more targeted queries" before searching. Google calls the same technique query fan-out: "a set of concurrent, related queries generated by the model."
  2. Candidates are retrieved from an index. That index is either the engine's own or a search provider's. OpenAI's crawler OAI-SearchBot exists to surface sites in ChatGPT search, and OpenAI also works with outside search providers. Perplexity runs its own index through PerplexityBot. Google's AI features draw on Google's own index.
  3. Pages are fetched and turned into text. Here the paths split sharply. Most AI fetchers request the raw HTML and do not run JavaScript; more on that below.
  4. Passages are selected. A model can only read so much text per answer. The pipeline picks the passages that look most relevant to the rewritten queries and discards the rest of the page.
  5. The answer is generated with citations. The model writes one response and attributes claims to a few sources.

Classic SEO mostly optimizes stage 2. GEO is about the other four, and stage 3 is where many modern sites quietly fail.

How AI search diverges from Googlebot

Rendering: most AI crawlers read raw HTML

Googlebot renders pages in a headless browser, so JavaScript-built content usually ends up indexed. Most AI crawlers don't render. In a December 2024 study across Vercel's network, Vercel and MERJ found that OpenAI's crawlers, Anthropic's ClaudeBot, Perplexity's bots, Meta's and ByteDance's all fetch pages without executing JavaScript. JavaScript files made up 11.5% of ChatGPT's fetches and 23.8% of Claude's, but neither crawler ran them. Only the crawlers built on Google's and Apple's infrastructure rendered JavaScript.

That matches Crawlable's own crawler registry. Of the 15 AI user agents the scanner checks, 13 read raw HTML only. A page that ranks well in Google can therefore be close to empty for ChatGPT, Claude and Perplexity. Why AI search engines can't read client-side rendered sites covers this in depth.

Separate bots for training and for answers

Googlebot does one job. AI companies split the work across several bots: one collects training data, one builds a search index, and one fetches a page at the moment a user asks about it. Blocking the training bot has no effect on whether you're cited. Blocking the search or user bots removes you from answers. Many robots.txt files get this backwards; our robots.txt guide for AI crawlers explains the split.

The unit that competes is a passage, not a page

A search engine ranks documents. An answer engine assembles an answer from excerpts. Your page doesn't need to be the best page overall. It needs to contain the clearest, most quotable passage for the specific sub-question the model is answering. A long page with one excellent paragraph can be cited. A strong page that buries its answer under a slow introduction may not be.

Token budgets: context is finite

A language model reads a limited amount of text, measured in tokens, for each answer. It also spreads that budget across several sources, so no single page gets read in full. The operators don't publish exact budgets for their retrieval pipelines, so treat the specifics as unknown. The principle still has practical consequences:

  • Answer early. A passage that states the answer in its first sentence or two has a better chance of making the cut than one that builds up to it.
  • Cut the noise around your content. Navigation, cookie text, repeated calls to action and huge inline scripts all compete with your prose when a page is turned into text.
  • One topic per section. A focused section under a descriptive heading is easier to select than a paragraph that covers three things.

Entities beat keywords

A generative engine is trying to understand things: your company, your product, the problem it solves, the people involved. It isn't matching strings. Make those relationships explicit in plain text. "Crawlable is an AI readability audit for websites" is a sentence a model can extract and reuse. "Unlock next-generation visibility" isn't.

Structured data such as Organization, Article and Product markup expresses the same facts in a machine-readable form. It's good practice, but don't overstate it: Google's May 2026 guide says "structured data isn't required for generative AI search." Clear writing does more of the work than markup does.

Why keyword stuffing fails in GEO

The GEO paper tested the oldest SEO trick directly. Rewriting sources to include more of the query's keywords reduced their position-adjusted word count, from 19.5 to 17.8 on the benchmark. When the same test ran against Perplexity itself, the score fell from 24.1 to 21.9. In the authors' words, keyword methods "offer little to no improvement on generative engine's responses."

The reason follows from the pipeline. Retrieval is increasingly semantic, so a page about the right thing is found whether or not it repeats the exact phrase. And the model choosing what to quote is looking for content that adds information. Repetition adds none.

What the research says works

The same paper found that three content changes reliably raised visibility:

MethodWhat it meansResult in the paper
Quotation additionQuoting a relevant, credible sourceAmong the strongest single methods
Statistics additionReplacing vague claims with specific numbersConsistently above baseline
Cite sourcesLinking the evidence behind your claimsStrongest when combined with other methods

The best methods improved position-adjusted word count by about 41% over the unmodified baseline. Two other findings matter more for small sites:

  • Lower-ranked pages gained the most. Adding citations raised visibility by 115% for pages that sat fifth in the underlying search results, while the top-ranked page lost 30% on average. Being retrieved at all matters. Once you're retrieved, the quality of your passage can beat a bigger site's authority.
  • Persuasive tone did nothing. Making text sound more authoritative without adding substance produced no significant gain.

Read these numbers with their limits in mind. The experiments used 2023-era models, mostly a simulated engine, and they measured share of the generated answer, not clicks or revenue. They tell you which way to push, not how much traffic you'll get.

What Google says about all this

For Google's own AI features, Google's position is blunt. Its guidance says "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." Its May 2026 optimization guide lists llms.txt files, content chunking and rewriting for AI as things you don't need.

That is consistent with everything above. Google renders JavaScript and draws on its own index, so good SEO already covers Google. GEO earns its name on the other engines, the ones that fetch raw HTML, run separate retrieval bots and pick passages from a different pool of sources.

A practical GEO checklist

Work through it in order. Each step depends on the one before it.

  1. Be fetchable. Allow the retrieval and user-action crawlers in robots.txt, check that your CDN or firewall isn't silently blocking them, and make sure key pages return 200.
  2. Be readable without JavaScript. The words that matter should be in the initial HTML response. Server-side rendering or static generation fixes this for most frameworks.
  3. Be extractable. Use semantic HTML: one h1, descriptive h2 sections, real lists and tables. Lead each section with its answer.
  4. Be citable. Make specific claims with numbers, dates and sources. Say plainly who you are and what you do.
  5. Be discoverable. Keep an accurate sitemap.xml. Consider an llms.txt file as a low-cost map for agents.
  6. Measure. Track AI referral traffic separately from search. Google SEO vs. AI search optimization shows how.

Steps 1 and 2 are where most sites lose, and they're the easiest to check: fetch your page the way a non-rendering crawler does and see what's left.