Skip to content
Crawlable

How to Make React Apps Crawlable for AI Search: SSR, Prerendering and Dynamic Rendering

Google can render React; ChatGPT, Claude and Perplexity crawlers can't. Code for Vite SSR, Next.js and Astro prerendering, dynamic rendering, and a curl check.

By Crawlable · · 5 min read

On this page

React apps get crawled. The question is what the crawler gets back. Googlebot fetches your HTML, queues the page, and later runs your JavaScript in headless Chromium. The crawlers behind ChatGPT search, Claude and Perplexity make one HTTP request and parse the response. If that response is Vite's default shell, they index an empty <div id="root"></div>.

This guide answers "can Google crawl React pages?", shows the AI-crawler view of the same page, and gives you the code for each fix: prerendering, server-side rendering, and dynamic rendering as a last resort.

Can Google crawl React pages?

Yes. Googlebot fetches the HTML first and then, in Google's words, "queues all pages with a 200 HTTP status code for rendering." A page "may stay on this queue for a few seconds, but it can take longer than that," after which "a headless Chromium renders the page and executes the JavaScript."

What still breaks on Google:

  • Links without an href. Googlebot follows <a href="/pricing">, not onClick={() => navigate('/pricing')} on a div.
  • Content behind interaction. Tabs, accordions and "load more" that fetch on click are never clicked.
  • Failed data calls. If your API rate-limits or blocks Googlebot, the render has nothing to show.
  • Titles set after hydration. react-helmet works for Google eventually, but the first-pass HTML carries Vite's default title.

What AI crawlers see on a React app

AI crawlers skip the render queue. A December 2024 study of crawler traffic across Vercel's network by Vercel and MERJ found that crawlers from OpenAI, Anthropic, Perplexity, Meta and ByteDance download JavaScript files but never execute them. Of the 15 AI user agents Crawlable tracks, 13 read raw HTML only.

Here's the whole of what they get from a fresh npm create vite@latest -- --template react-ts build:

<!doctype html>
<html lang="en">
  <head>
    <meta charset="UTF-8" />
    <meta name="viewport" content="width=device-width, initial-scale=1.0" />
    <title>Vite + React + TS</title>
    <script type="module" crossorigin src="/assets/index-a1b2c3.js"></script>
  </head>
  <body>
    <div id="root"></div>
  </body>
</html>

No headline, no body copy, no internal links, and a title that names the bundler. Every fix below has one goal: put the rendered markup in that response.

Verify the raw HTML with curl

Before changing anything, look at what a non-rendering crawler receives. The short form:

curl -sS -A "PerplexityBot" https://yoursite.com/pricing | head -c 2000

The full user agent Perplexity documents, saved to a file you can inspect:

UA='Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)'
curl -sS -A "$UA" -o raw.html -w "%{http_code} %{size_download} bytes\n" https://yoursite.com/pricing

grep -o '<title>[^<]*</title>' raw.html
grep -c '<h1' raw.html
grep -o 'id="root">[^<]\{0,80\}' raw.html   # empty after the > means client-rendered

# The words a non-rendering crawler can read: scripts, styles and tags stripped
python3 -c "import re;h=open('raw.html').read();h=re.sub(r'(?is)<(script|style)[^>]*>.*?</\1>',' ',h);print(' '.join(re.sub(r'<[^>]+>',' ',h).split())[:600])"

Two things to read into the result:

  • An empty #root and no <h1>: the page is client-rendered. AI crawlers see nothing.
  • A 403 or a challenge page: your CDN or WAF is blocking requests that claim to be PerplexityBot. The real bot comes from Perplexity's published IP ranges, so a verified-bot rule may treat it differently from your curl. Check the firewall's logs before concluding either way. The PerplexityBot page covers verification.

Fix 1: prerender (static site generation)

If you know your routes at build time (marketing pages, docs, a blog, product pages from a CMS), generate HTML for each one at build time. It's the cheapest fix to run: the output is static files on a CDN, and every crawler gets complete HTML.

React Router v7 on Vite

React Router's framework mode prerenders routes at build time, and with ssr: false the build is still a static deployment with no server:

// react-router.config.ts
import type { Config } from '@react-router/dev/config';

export default {
  ssr: false,
  async prerender() {
    // Static paths, plus any you can list from your CMS or filesystem.
    const posts = await fetch('https://cms.example.com/api/posts').then((r) => r.json());
    return ['/', '/pricing', '/docs', ...posts.map((p: { slug: string }) => `/blog/${p.slug}`)];
  },
} satisfies Config;

Each route's loader runs during the build, so data fetched there is in the HTML. Data fetched in useEffect is not: effects don't run during prerendering.

// app/routes/pricing.tsx
import type { Route } from './+types/pricing';

export async function loader() {
  const plans = await fetch('https://cms.example.com/api/plans').then((r) => r.json());
  return { plans };
}

export function meta() {
  return [
    { title: 'Pricing | Acme' },
    { name: 'description', content: 'Plans for teams of every size.' },
  ];
}

export default function Pricing({ loaderData }: Route.ComponentProps) {
  return (
    <main>
      <h1>Pricing</h1>
      {loaderData.plans.map((plan: { id: string; name: string; summary: string }) => (
        <section key={plan.id}>
          <h2>{plan.name}</h2>
          <p>{plan.summary}</p>
        </section>
      ))}
    </main>
  );
}

Next.js: static params and server components

In the App Router, server components render to HTML by default. generateStaticParams turns a dynamic route into static pages at build time:

// app/docs/[slug]/page.tsx
import type { Metadata } from 'next';
import { getDoc, listDocs } from '@/lib/docs';

export async function generateStaticParams() {
  return (await listDocs()).map((doc) => ({ slug: doc.slug }));
}

export async function generateMetadata({ params }: { params: Promise<{ slug: string }> }): Promise<Metadata> {
  const doc = await getDoc((await params).slug);
  return { title: doc.title, description: doc.summary };
}

export default async function DocPage({ params }: { params: Promise<{ slug: string }> }) {
  const doc = await getDoc((await params).slug);
  return (
    <article>
      <h1>{doc.title}</h1>
      <div dangerouslySetInnerHTML={{ __html: doc.html }} />
    </article>
  );
}

The trap in Next.js is the same as in Vite: a 'use client' component that fetches its content in useEffect ships an empty container even on a statically generated page. Fetch in the server component and pass the data down as props. The Next.js platform guide covers the other Next.js-specific checks.

Astro: static by default, React where it's interactive

Astro renders pages to HTML at build time and ships no JavaScript unless you opt a component in. Your existing React components render to HTML; client:* directives hydrate only the interactive ones:

---
// src/pages/docs/[slug].astro
import { getCollection, render } from 'astro:content';
import Layout from '../../layouts/Layout.astro';
import SearchBox from '../../components/SearchBox.tsx';

export async function getStaticPaths() {
  const docs = await getCollection('docs');
  return docs.map((doc) => ({ params: { slug: doc.id }, props: { doc } }));
}

const { doc } = Astro.props;
const { Content } = await render(doc);
---
<Layout title={doc.data.title} description={doc.data.description}>
  <h1>{doc.data.title}</h1>
  <Content />
  <SearchBox client:idle />
</Layout>

Next.js vs Astro for crawlable React

Next.js (App Router)Astro
Default output for a pageServer-rendered HTML, static when nothing is request-specificStatic HTML
JavaScript shipped by defaultThe React runtime plus client componentsNone; only components with a client:* directive
React componentsServer and client componentsRendered to HTML; hydrated per island
Request-time pagesBuilt in: per-request rendering, ISRWith an SSR adapter, per page via export const prerender = false
Best fitAn app with lots of interactive, authenticated UI and some public pagesContent-heavy sites with islands of interactivity

Both produce complete HTML for crawlers. The choice is about how much client-side app you have, not crawlability.

Fix 2: server-side rendering on Vite

When a page depends on the request (search results, per-user content, inventory that changes by the minute) render it on the server. Vite's SSR setup needs three pieces: a server entry, a client entry that hydrates, and a server that stitches them together.

// src/entry-server.tsx
import { renderToString } from 'react-dom/server';
import { StaticRouter } from 'react-router';
import { App } from './App';
import { loadRouteData, type RouteData } from './data';

export async function render(url: string): Promise<{ html: string; data: RouteData }> {
  // Load data BEFORE rendering. renderToString does not wait for effects or promises.
  const data = await loadRouteData(url);
  const html = renderToString(
    <StaticRouter location={url}>
      <App data={data} />
    </StaticRouter>,
  );
  return { html, data };
}
// src/entry-client.tsx
import { hydrateRoot } from 'react-dom/client';
import { BrowserRouter } from 'react-router';
import { App } from './App';
import type { RouteData } from './data';

declare global {
  interface Window {
    __DATA__: RouteData;
  }
}

hydrateRoot(
  document.getElementById('root')!,
  <BrowserRouter>
    <App data={window.__DATA__} />
  </BrowserRouter>,
);
// server.js (production: serves the client build, renders with the SSR build)
import fs from 'node:fs/promises';
import express from 'express';

const template = await fs.readFile('./dist/client/index.html', 'utf-8');
const { render } = await import('./dist/server/entry-server.js');

const app = express();
app.use(express.static('./dist/client', { index: false }));

app.use(async (req, res, next) => {
  try {
    const { html, data } = await render(req.originalUrl);
    // Escape "<" so data containing "</script>" cannot end the tag early.
    const state = JSON.stringify(data).replace(/</g, '\\u003c');
    res
      .status(200)
      .set('Content-Type', 'text/html')
      .end(
        template
          .replace('<!--ssr-outlet-->', html)
          .replace('<!--ssr-state-->', `<script>window.__DATA__=${state}</script>`),
      );
  } catch (error) {
    next(error);
  }
});

app.listen(3000);

index.html gets the two placeholders: <div id="root"><!--ssr-outlet--></div> and <!--ssr-state--> before your module script. Build both bundles:

vite build --outDir dist/client
vite build --outDir dist/server --ssr src/entry-server.tsx

If you're starting fresh, React Router's framework mode with ssr: true (the default) gives you the same result without the hand-written server.

Fix 3: dynamic rendering, the workaround

Dynamic rendering serves crawlers a prerendered snapshot and users the normal client-side app. Google describes it as "a workaround and not a recommended solution" and recommends server-side rendering, static rendering or hydration instead. Use it when you can't change the app: a legacy SPA, a vendor build, a migration that's months away.

Step one: snapshot each route with a headless browser after the build.

// scripts/snapshot.ts (run against `vite preview`, which serves on :4173)
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { chromium } from 'playwright';

const ORIGIN = 'http://localhost:4173';
const ROUTES = ['/', '/pricing', '/docs/getting-started'];

const browser = await chromium.launch();
const page = await browser.newPage();
for (const route of ROUTES) {
  await page.goto(ORIGIN + route, { waitUntil: 'networkidle' });
  const dir = path.join('dist-prerendered', route);
  await mkdir(dir, { recursive: true });
  await writeFile(path.join(dir, 'index.html'), await page.content());
}
await browser.close();

Step two: route non-rendering AI crawlers to the snapshots. This nginx config is generated from the crawler registry Crawlable scans against, so the user-agent list matches what the scanner checks:

# Generated from the crawler registry Crawlable scans against.
map $http_user_agent $prerender_bot {
  default 0;
  "~*(GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Amazonbot|Bytespider|meta-externalagent|cohere-ai|MistralAI-User)" 1;
}

server {
  listen 443 ssl;
  server_name example.com;

  location / {
    root /var/www/spa/dist;                # the client-side build
    if ($prerender_bot) {
      root /var/www/spa/dist-prerendered;  # one index.html per route
    }
    try_files $uri $uri/index.html /index.html;
  }
}

Keep two rules. The snapshot must carry the same content users see; Google's documentation doesn't treat dynamic rendering as cloaking on that condition. And re-snapshot on every deploy, or crawlers read last month's prices.

Which fix to use

  • Routes known at build time: prerender. React Router prerender(), Next.js generateStaticParams, or Astro.
  • Request-dependent pages: server-side rendering.
  • Can't touch the app: dynamic rendering, with snapshots rebuilt per deploy.

Whichever you pick, re-run the curl check from earlier against production. The client-side rendering explainer covers why the gap exists, what GPTBot sees on JavaScript sites walks through a hydration failure in detail, and the React platform page lists the React-specific checks.

Frequently asked questions

Can Google crawl React pages?
Yes. Googlebot queues pages for rendering and runs them in an evergreen headless Chromium, so content a React app renders in the browser usually gets indexed. Rendering happens in a second pass that can lag the crawl, and content behind clicks, failed API calls or links without a real href can still be missed.
Can ChatGPT, Claude and Perplexity crawl a React website?
They can fetch it, but their crawlers don't run JavaScript. Vercel and MERJ found that crawlers from OpenAI, Anthropic, Perplexity, Meta and ByteDance download JavaScript files without executing them. A client-rendered React app gives them an empty root div, so the content has to be in the HTML the server sends.
What is the best way to prerender a React app?
If the routes are known at build time, generate static HTML for them: React Router's prerender option for a Vite app, generateStaticParams in Next.js, or Astro's static output. If pages depend on the request, render them on the server. Dynamic rendering (serving bots a snapshot) is the fallback when neither is possible.
Is dynamic rendering cloaking?
Google's documentation says dynamic rendering is generally not treated as cloaking as long as the snapshot carries the same content users see. Serving crawlers different content from users is cloaking. Google also calls dynamic rendering a workaround and recommends server-side rendering, static rendering or hydration instead.
How do I check what an AI crawler sees on my React site?
Request the page with curl and a crawler's user agent, for example curl -A PerplexityBot, without running JavaScript, and look for your headline and body text in the response. If the HTML holds only an empty root div and script tags, AI crawlers see an empty page.