Skip to content
Crawlable

CRAWLABLE RESEARCH ·

The 2026 AI Search Readability Index: 50 Top Tech & DTC Brands Audited

We ran 50 leading startup, SaaS and e-commerce sites through Crawlable's production audit. None scored an A. 78% don't name a single AI crawler in robots.txt, and 70% serve no llms.txt that follows the spec.

Key findings

  • 51.5/100

    Average AI readability score

    Median 55. No site reached grade A (90+); 43 of 46 scored below 70.

  • 70%

    No spec-compliant llms.txt

    11 serve none; 21 serve one our validator rejects, 8 of them Shopify's auto-generated default.

  • 30%

    Ship empty JavaScript shells

    14 of 46 sites served at least one audited page as an empty shell, with the text arriving only after hydration.

  • 78%

    Name no AI crawler in robots.txt

    None block OAI-SearchBot or PerplexityBot. The gap is silence, not blocking.

Scores by cohort

Average total score out of 100 for the sites we could measure in each cohort, with the four pillars that make it up.

YC & high-growth startups

16 of 17 measured

43.1/100

Major SaaS & tech companies

16 of 17 measured

59.2/100

E-commerce & DTC brands

14 of 16 measured

52.4/100
Average pillar points by cohort
CohortRaw-HTML readability /35Schema depth /25Context architecture (llms.txt) /20Crawler governance /20Total /100
YC & high-growth startups24.115.62.80.643.1
Major SaaS & tech companies28.117.810.34.459.2
E-commerce & DTC brands26.4205.70.752.4
All 46 sites26.217.76.3251.5

Where the points are lost

Average points earned on each pillar across all 46 measured sites, against the maximum available.

Raw-HTML readability

of 35

26.2/35

Schema depth

of 25

17.7/25

Context architecture (llms.txt)

of 20

6.3/20

Crawler governance

of 20

2/20

Key takeaways

1. Crawler rules are silent, not hostile

Not one measured site blocks OAI-SearchBot or PerplexityBot from its home page. 2 block another answer-engine crawler (Amazonbot). The bigger pattern is that 36 of 46 sites name no AI crawler at all, so every bot falls back to the generic * rules. Unless those rules block them, that lets them in, but it states no intent: there is no separate decision for search crawlers and training crawlers, and Crawlable's governance pillar scores that 0 of 20. Naming retrieval bots and training bots in separate groups takes a few lines. The platform guides have the code.

2. Most sites have an llms.txt; most of those fail the spec

14 sites serve an llms.txt that passes Crawlable's validator: one # heading, a > summary, ## sections, and absolute links with a description after each. 35 of 46 serve a file at all; 21 serve a file that misses at least one of those, most often the summary line or the link descriptions, and 11 serve none. 8 of the files that fail are Shopify's auto-generated "Agent Instructions", which every store has served by default since May 2026: it documents checkout endpoints for shopping agents rather than mapping the catalog. A store can replace it with a theme template; the Shopify guide shows how.

3. Empty JavaScript shells: 14 of 46 sites

10 of 46 sites earned the full 35 readability points. Of the 36 that did not, 22 had no empty page at all and lost points to short or markup-heavy pages. But 14 of 46 sites served at least one audited page as an empty JavaScript shell, where the text arrives only after the browser runs the app. Among the 26 sites detected as Next.js, 11 had at least one such page. The extreme case was v0.app, where 37 of 40 audited pages were shells. The Next.js guide shows the usual cause, content fetched after the page mounts, and the fix. Shells cost more than their own pages: the rendering gate capped 4 sites at 64.

4. Schema depth: 27 of 46 sites complete

27 of 46 sites carry a complete rich entity, such as a SoftwareApplication with an offer or an article with its author, and earn all 25 schema points. 12 of 46 sites have nothing beyond Organization or WebSite markup, or no valid JSON-LD at all.

Every site, every pillar

Scores as measured on 11 October 2026. Sites change; a re-scan today may differ. Rows marked "Not measurable" refused our crawler outright and carry no score.

AI readability results for every domain in the study
DomainCohortRaw HTML /35Schema /25llms.txtCrawler rulesGEO score
modal.comStartup3525Not to spec1 AI bots named65 D
runwayml.comaudited at runway.comStartup2525Well-formedNo AI bot named65 D
supabase.comStartup2525Not to spec4 AI bots named65 D
mistral.aiStartup2515Well-formedNo AI bot named55 D
resend.comStartup2525Not to specNo AI bot named55 D
cursor.comStartup2525Not to specNo AI bot named50 D
dub.coStartup2525Not to specNo AI bot named50 D
elevenlabs.ioStartup2525Not to specNo AI bot named50 D
posthog.comStartup2525Not to specNo AI bot named50 D
replicate.comStartup2525Not to specNo AI bot named50 D
raycast.comStartup350MissingNo AI bot named35 F
scale.comStartup255MissingNo AI bot named30 F
cal.comStartup250Not to specNo AI bot named25 F
linear.appStartup250Not to specNo AI bot named25 F
midjourney.comStartup150MissingBlocks Amazonbot15 F
v0.devaudited at v0.appStartup05MissingNo AI bot named5 F
perplexity.aiStartupNot measurable: refused our crawler (HTTP 403)
hubspot.comSaaS3525Well-formedNo AI bot named75 C
clickup.comSaaS3515Well-formedNo AI bot named70 C
figma.comSaaS2525Missing10 AI bots named70 C
intercom.comSaaS2525Well-formedNo AI bot named65 D
zapier.comSaaS2525Not to spec8 AI bots named65 D
asana.comSaaS2525Well-formedNo AI bot named64 D
monday.comSaaS2525Well-formed6 AI bots named64 D
vercel.comSaaS2525Well-formedNo AI bot named64 D
loom.comSaaS355Not to spec2 AI bots named60 D
notion.soaudited at www.notion.comSaaS2515Well-formedBlocks Amazonbot55 D
stripe.comSaaS355Well-formedNo AI bot named55 D
airtable.comSaaS2525MissingNo AI bot named50 D
miro.comSaaS2525MissingNo AI bot named50 D
shopify.comSaaS255Well-formedNo AI bot named50 D
slack.comSaaS350Well-formedNo AI bot named50 D
webflow.comSaaS2515MissingNo AI bot named40 F
canva.comSaaSNot measurable: refused our crawler (HTTP 403)
allbirds.comE-commerce3525Not to specNo AI bot named65 D
awaytravel.comE-commerce2525Not to spec11 AI bots named65 D
glossier.comE-commerce3525Not to specNo AI bot named65 D
magicspoon.comE-commerce3525Not to specNo AI bot named65 D
warbyparker.comE-commerce2525Well-formed2 AI bots named64 D
casper.comE-commerce2525Not to specNo AI bot named55 D
outdoorvoices.comE-commerce2525Not to specNo AI bot named55 D
rothys.comE-commerce2525Not to specNo AI bot named55 D
brooklinen.comE-commerce2515Not to specNo AI bot named45 F
untuckit.comE-commerce2515Not to specNo AI bot named45 F
vuoriclothing.comE-commerce255Well-formedNo AI bot named45 F
drinkpoppi.comE-commerce2515MissingNo AI bot named40 F
skims.comE-commerce1525MissingNo AI bot named40 F
gymshark.comE-commerce255MissingNo AI bot named30 F
bombas.comE-commerceNot measurable: refused our crawler (HTTP 429)
liquidiv.comE-commerceNot measurable: refused our crawler (HTTP 503)

How we measured

  • Every domain was run through Crawlable's production audit engine (commit 5857303, scoring v2), the same code that runs a customer audit, on 11 October 2026. Nothing was scored by hand.
  • Up to 40 pages per site, chosen from the sitemap or the home page's links, within a 45-second budget, 5 sites at a time. Each page was fetched once, as raw HTML, with no JavaScript run.
  • Requests identified themselves as CrawlableBot/1.0 (+https://usecrawlable.com/bot; AI-readability auditor) and came from a cloud server. 4 sites (perplexity.ai, canva.com, liquidiv.com, bombas.com) refused that crawler with a bot challenge or an error page, so they are listed but excluded from every average. That says nothing about how those sites treat verified AI crawlers, which come from their operators' own networks.
  • The score is the sum of four pillars: raw-HTML readability (35), schema depth (25), llms.txt (20) and crawler governance (20), capped at 64 when a page arrives as an empty JavaScript shell. Robots rules were evaluated for the home page path.
  • Two domains in the original list were corrected (bombas.com, rothys.com). The bare domain redirected our crawler to us.checkout.gymshark.com, so the storefront at www.gymshark.com was audited.
  • The full per-domain data, including the validator's reasons for every llms.txt it rejected, is in the CSV and JSON.