What GPTBot Actually Sees When Your Site Relies on Client-Side JavaScript
AI crawlers like GPTBot and OAI-SearchBot fetch raw HTML without executing client-side JS. Here is what they actually extract from SPAs.
By Shahid Saleem · · 6 min read
On this page
Open your pricing page in a browser and you see plans, prices and a feature list. Request the same URL the way GPTBot does and you may get a loading spinner and a list of script tags. Both are the same page, at different moments: the browser waits for your JavaScript to run, and the crawler doesn't.
This post shows exactly what that difference looks like on the wire, how to check your own pages in one command, and the three changes that fix it.
Rendering costs compute; answers cost time
Googlebot reads JavaScript-heavy sites because Google built a second pass for them. Its crawler fetches the HTML first, then, in Google's words, "queues all pages with a 200 HTTP status code for rendering." The page "may stay on this queue for a few seconds, but it can take longer than that." When resources allow, "a headless Chromium renders the page and executes the JavaScript." Nobody is waiting on that queue, so Google can afford to be patient.
AI crawlers skip the second pass. A December 2024 study of crawler traffic by Vercel and MERJ found that OpenAI's crawlers, GPTBot, OAI-SearchBot and ChatGPT-User, download JavaScript files but don't execute them. Each one makes an HTTP request, takes the HTML that comes back, and works with that.
The engineering reasons are straightforward, though no operator publishes them as policy:
- Compute. OpenAI's crawler documentation describes GPTBot as collecting content that may be used for training, and OAI-SearchBot as surfacing sites in ChatGPT search. Both crawl at web scale. A headless browser per URL costs far more CPU and memory than a plain HTTP fetch.
- Time. ChatGPT-User fetches a page while a person waits for an answer. OpenAI doesn't publish a timeout, but a fetch inside a chat response can't wait for an app bundle to download, boot, call its APIs and settle.
The result is the same for all three: the first HTML response is the whole page, as far as they're concerned.
The hydration void
Here is a pricing component written the way many React apps are. It's a client component that fetches its data after it mounts:
'use client';
import { useEffect, useState } from 'react';
type Plan = { name: string; price: string; features: string[] };
export default function Pricing() {
const [plans, setPlans] = useState<Plan[] | null>(null);
useEffect(() => {
fetch('/api/plans')
.then((res) => res.json())
.then(setPlans);
}, []);
if (!plans) return <div className="spinner">Loading plans...</div>;
return (
<section>
<h2>Pricing</h2>
{plans.map((plan) => (
<article key={plan.name}>
<h3>{plan.name}</h3>
<p>{plan.price}</p>
<ul>
{plan.features.map((feature) => (
<li key={feature}>{feature}</li>
))}
</ul>
</article>
))}
</section>
);
}
In Next.js, 'use client' doesn't stop the server from rendering this component. It renders the initial state on the server, and useEffect never runs there, so the initial state is the spinner. This is the HTML that goes over the wire:
<!doctype html>
<html lang="en">
<head>
<title>Pricing | Acme</title>
<script src="/_next/static/chunks/app/pricing/page-3f2a1c.js" async></script>
</head>
<body>
<main>
<div class="spinner">Loading plans...</div>
</main>
<script>self.__next_f.push([1,"..."])</script>
</body>
</html>
Strip the markup and scripts, and the text a crawler can extract is "Pricing | Acme" and "Loading plans...". There are no plan names, prices or features in it. A Vite or Create React App single-page app is emptier still: usually a div id="root" and nothing inside it.
What happens downstream
An answer engine that retrieves pages turns them into text, splits the text into passages, and hands the most relevant ones to the model as context. The internals differ between products, but every version can only work with text it received. From this page it received two strings, and neither contains a fact.
So when someone asks what your product costs, the model has nothing from your site to quote. It answers from whatever else it found, such as a review site, a competitor's comparison page or an outdated snapshot, or it leaves you out. Your page exists, ranks on Google, and still contributes nothing to the answer.
Test it from a terminal
Request the page with the crawler's user agent and drop the script and style lines:
curl -s -A "GPTBot" https://yourdomain.com | grep -v "<script" | grep -v "<style"
Read what's left. If you see your headline, prices and body copy, the content is in the HTML. If you see a spinner, an empty div or nothing, a non-rendering crawler sees the same.
Two cautions before you trust the result:
- Missing text isn't proof.
grep -vremoves whole lines, and frameworks often put your content on the same line as a script tag. Run against this post, which is fully server-rendered, the command drops theh1and the opening paragraph. Before concluding something is absent, search for it directly:curl -s -A "GPTBot" https://yourdomain.com/pricing | grep -c "Pro plan". Any count above zero means that text is in the raw HTML. - Check that you got the real page. Some CDNs and firewalls block AI user agents or answer them with a challenge page. A 403 or a CAPTCHA tells you about your firewall rules, not your rendering.
For a word count across the whole page, the client-side rendering guide has a longer command that strips tags and counts what remains.
Three rules that fix it
1. Default to Server Components
Fetch on the server and render the result, so the facts are in the first bytes of the response:
// app/pricing/page.tsx: a Server Component, with no 'use client'
import { getPlans } from '@/lib/plans';
export default async function PricingPage() {
const plans = await getPlans();
return (
<section>
<h2>Pricing</h2>
{plans.map((plan) => (
<article key={plan.name}>
<h3>{plan.name}</h3>
<p>{plan.price}</p>
</article>
))}
</section>
);
}
Keep 'use client' for the parts that need the browser, like a billing-period toggle, and pass them data the server already fetched. A rule of thumb: if a sentence would answer a buyer's question, it belongs in the server render. The Next.js and React guides cover the framework-specific details.
The same rule applies outside the App Router. In the Pages Router, move the fetch into getStaticProps or getServerSideProps so the props arrive with the HTML. In Nuxt, SvelteKit and Angular, use the framework's server data loading rather than a fetch in a mount hook. For pages whose facts change rarely, such as pricing, docs and comparison pages, static generation at build time is the simplest option: the HTML is complete before any request arrives, and it costs nothing per visit.
2. Publish /llms.txt for the facts that matter
An llms.txt file is a Markdown summary at the root of your site: what you sell, who it's for, what it costs, and links to the pages that explain more. It has no layout and no scripts, so an agent that reads it gets dense facts in a few hundred words.
Be clear about what it is. It's a proposed standard aimed at assistants and agents reading your site to answer a question. Google says you "don't need to create new machine readable files" to appear in its search, and no answer engine has said it ranks sites by it. It's cheap to add and doesn't replace rule 1. The llms.txt guide walks through the format.
3. Treat search crawlers and training crawlers differently
The bots in your logs do different jobs. OAI-SearchBot and PerplexityBot build the indexes that answer engines cite from. GPTBot collects training data, and Common Crawl's CCBot builds an open archive that many training sets draw on. Neither of those two feeds a live answer.
That gives you a clean split in robots.txt:
# Answer engines: allow them, or you can't be cited
User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /
# Training crawlers: your choice. Blocking them doesn't remove you from answers.
User-agent: GPTBot
User-agent: CCBot
Disallow: /
Whether to block training crawlers is a business decision. Blocking a search crawler costs you citations. The robots.txt guide has a complete file covering every major crawler.
None of these rules works on its own. Allowing OAI-SearchBot does nothing if the page it fetches is a spinner, and the best llms.txt can't stand in for pages that say nothing until JavaScript runs. Start with rule 1.
Keep reading
- Why AI Search Engines Can’t Read Client-Side Rendered (CSR) Sites
Most AI crawlers fetch raw HTML and never run your JavaScript. Here's what they actually see on a client-rendered site, how to test it, and how to fix it.
- How to Create and Optimize an llms.txt File for Your Website
The llms.txt format, where the file goes, how to keep it token-efficient, working examples for Next.js and static sites, and an honest account of who uses it.
- How to Configure robots.txt for AI Crawlers (Without Compromising Security)
Allow the AI crawlers that get you cited, opt out of the ones that only train models, and avoid treating robots.txt as a security control. Includes a ready-made file.