CRAWLABLE RESEARCH ·
The 2026 AI Search Readability Index: 50 Top Tech & DTC Brands Audited
We ran 50 leading startup, SaaS and e-commerce sites through Crawlable's production audit. None scored an A. 78% don't name a single AI crawler in robots.txt, and 70% serve no llms.txt that follows the spec.
Key findings
51.5/100
Average AI readability score
Median 55. No site reached grade A (90+); 43 of 46 scored below 70.
70%
No spec-compliant llms.txt
11 serve none; 21 serve one our validator rejects, 8 of them Shopify's auto-generated default.
30%
Ship empty JavaScript shells
14 of 46 sites served at least one audited page as an empty shell, with the text arriving only after hydration.
78%
Name no AI crawler in robots.txt
None block OAI-SearchBot or PerplexityBot. The gap is silence, not blocking.
Scores by cohort
Average total score out of 100 for the sites we could measure in each cohort, with the four pillars that make it up.
YC & high-growth startups
16 of 17 measured
Major SaaS & tech companies
16 of 17 measured
E-commerce & DTC brands
14 of 16 measured
| Cohort | Raw-HTML readability /35 | Schema depth /25 | Context architecture (llms.txt) /20 | Crawler governance /20 | Total /100 |
|---|---|---|---|---|---|
| YC & high-growth startups | 24.1 | 15.6 | 2.8 | 0.6 | 43.1 |
| Major SaaS & tech companies | 28.1 | 17.8 | 10.3 | 4.4 | 59.2 |
| E-commerce & DTC brands | 26.4 | 20 | 5.7 | 0.7 | 52.4 |
| All 46 sites | 26.2 | 17.7 | 6.3 | 2 | 51.5 |
Where the points are lost
Average points earned on each pillar across all 46 measured sites, against the maximum available.
Raw-HTML readability
of 35
Schema depth
of 25
Context architecture (llms.txt)
of 20
Crawler governance
of 20
Key takeaways
1. Crawler rules are silent, not hostile
Not one measured site blocks OAI-SearchBot or PerplexityBot from its home page. 2 block another answer-engine crawler (Amazonbot). The bigger pattern is that 36 of 46 sites name no AI crawler at all, so every bot falls back to the generic * rules. Unless those rules block them, that lets them in, but it states no intent: there is no separate decision for search crawlers and training crawlers, and Crawlable's governance pillar scores that 0 of 20. Naming retrieval bots and training bots in separate groups takes a few lines. The platform guides have the code.
2. Most sites have an llms.txt; most of those fail the spec
14 sites serve an llms.txt that passes Crawlable's validator: one # heading, a > summary, ## sections, and absolute links with a description after each. 35 of 46 serve a file at all; 21 serve a file that misses at least one of those, most often the summary line or the link descriptions, and 11 serve none. 8 of the files that fail are Shopify's auto-generated "Agent Instructions", which every store has served by default since May 2026: it documents checkout endpoints for shopping agents rather than mapping the catalog. A store can replace it with a theme template; the Shopify guide shows how.
3. Empty JavaScript shells: 14 of 46 sites
10 of 46 sites earned the full 35 readability points. Of the 36 that did not, 22 had no empty page at all and lost points to short or markup-heavy pages. But 14 of 46 sites served at least one audited page as an empty JavaScript shell, where the text arrives only after the browser runs the app. Among the 26 sites detected as Next.js, 11 had at least one such page. The extreme case was v0.app, where 37 of 40 audited pages were shells. The Next.js guide shows the usual cause, content fetched after the page mounts, and the fix. Shells cost more than their own pages: the rendering gate capped 4 sites at 64.
4. Schema depth: 27 of 46 sites complete
27 of 46 sites carry a complete rich entity, such as a SoftwareApplication with an offer or an article with its author, and earn all 25 schema points. 12 of 46 sites have nothing beyond Organization or WebSite markup, or no valid JSON-LD at all.
Every site, every pillar
Scores as measured on 11 October 2026. Sites change; a re-scan today may differ. Rows marked "Not measurable" refused our crawler outright and carry no score.
| Domain | Cohort | Raw HTML /35 | Schema /25 | llms.txt | Crawler rules | GEO score |
|---|---|---|---|---|---|---|
| modal.com | Startup | 35 | 25 | Not to spec | 1 AI bots named | 65 D |
| runwayml.comaudited at runway.com | Startup | 25 | 25 | Well-formed | No AI bot named | 65 D |
| supabase.com | Startup | 25 | 25 | Not to spec | 4 AI bots named | 65 D |
| mistral.ai | Startup | 25 | 15 | Well-formed | No AI bot named | 55 D |
| resend.com | Startup | 25 | 25 | Not to spec | No AI bot named | 55 D |
| cursor.com | Startup | 25 | 25 | Not to spec | No AI bot named | 50 D |
| dub.co | Startup | 25 | 25 | Not to spec | No AI bot named | 50 D |
| elevenlabs.io | Startup | 25 | 25 | Not to spec | No AI bot named | 50 D |
| posthog.com | Startup | 25 | 25 | Not to spec | No AI bot named | 50 D |
| replicate.com | Startup | 25 | 25 | Not to spec | No AI bot named | 50 D |
| raycast.com | Startup | 35 | 0 | Missing | No AI bot named | 35 F |
| scale.com | Startup | 25 | 5 | Missing | No AI bot named | 30 F |
| cal.com | Startup | 25 | 0 | Not to spec | No AI bot named | 25 F |
| linear.app | Startup | 25 | 0 | Not to spec | No AI bot named | 25 F |
| midjourney.com | Startup | 15 | 0 | Missing | Blocks Amazonbot | 15 F |
| v0.devaudited at v0.app | Startup | 0 | 5 | Missing | No AI bot named | 5 F |
| perplexity.ai | Startup | Not measurable: refused our crawler (HTTP 403) | ||||
| hubspot.com | SaaS | 35 | 25 | Well-formed | No AI bot named | 75 C |
| clickup.com | SaaS | 35 | 15 | Well-formed | No AI bot named | 70 C |
| figma.com | SaaS | 25 | 25 | Missing | 10 AI bots named | 70 C |
| intercom.com | SaaS | 25 | 25 | Well-formed | No AI bot named | 65 D |
| zapier.com | SaaS | 25 | 25 | Not to spec | 8 AI bots named | 65 D |
| asana.com | SaaS | 25 | 25 | Well-formed | No AI bot named | 64 D |
| monday.com | SaaS | 25 | 25 | Well-formed | 6 AI bots named | 64 D |
| vercel.com | SaaS | 25 | 25 | Well-formed | No AI bot named | 64 D |
| loom.com | SaaS | 35 | 5 | Not to spec | 2 AI bots named | 60 D |
| notion.soaudited at www.notion.com | SaaS | 25 | 15 | Well-formed | Blocks Amazonbot | 55 D |
| stripe.com | SaaS | 35 | 5 | Well-formed | No AI bot named | 55 D |
| airtable.com | SaaS | 25 | 25 | Missing | No AI bot named | 50 D |
| miro.com | SaaS | 25 | 25 | Missing | No AI bot named | 50 D |
| shopify.com | SaaS | 25 | 5 | Well-formed | No AI bot named | 50 D |
| slack.com | SaaS | 35 | 0 | Well-formed | No AI bot named | 50 D |
| webflow.com | SaaS | 25 | 15 | Missing | No AI bot named | 40 F |
| canva.com | SaaS | Not measurable: refused our crawler (HTTP 403) | ||||
| allbirds.com | E-commerce | 35 | 25 | Not to spec | No AI bot named | 65 D |
| awaytravel.com | E-commerce | 25 | 25 | Not to spec | 11 AI bots named | 65 D |
| glossier.com | E-commerce | 35 | 25 | Not to spec | No AI bot named | 65 D |
| magicspoon.com | E-commerce | 35 | 25 | Not to spec | No AI bot named | 65 D |
| warbyparker.com | E-commerce | 25 | 25 | Well-formed | 2 AI bots named | 64 D |
| casper.com | E-commerce | 25 | 25 | Not to spec | No AI bot named | 55 D |
| outdoorvoices.com | E-commerce | 25 | 25 | Not to spec | No AI bot named | 55 D |
| rothys.com | E-commerce | 25 | 25 | Not to spec | No AI bot named | 55 D |
| brooklinen.com | E-commerce | 25 | 15 | Not to spec | No AI bot named | 45 F |
| untuckit.com | E-commerce | 25 | 15 | Not to spec | No AI bot named | 45 F |
| vuoriclothing.com | E-commerce | 25 | 5 | Well-formed | No AI bot named | 45 F |
| drinkpoppi.com | E-commerce | 25 | 15 | Missing | No AI bot named | 40 F |
| skims.com | E-commerce | 15 | 25 | Missing | No AI bot named | 40 F |
| gymshark.com | E-commerce | 25 | 5 | Missing | No AI bot named | 30 F |
| bombas.com | E-commerce | Not measurable: refused our crawler (HTTP 429) | ||||
| liquidiv.com | E-commerce | Not measurable: refused our crawler (HTTP 503) | ||||
How we measured
- Every domain was run through Crawlable's production audit engine (commit
5857303, scoring v2), the same code that runs a customer audit, on 11 October 2026. Nothing was scored by hand. - Up to 40 pages per site, chosen from the sitemap or the home page's links, within a 45-second budget, 5 sites at a time. Each page was fetched once, as raw HTML, with no JavaScript run.
- Requests identified themselves as
CrawlableBot/1.0 (+https://usecrawlable.com/bot; AI-readability auditor)and came from a cloud server. 4 sites (perplexity.ai, canva.com, liquidiv.com, bombas.com) refused that crawler with a bot challenge or an error page, so they are listed but excluded from every average. That says nothing about how those sites treat verified AI crawlers, which come from their operators' own networks. - The score is the sum of four pillars: raw-HTML readability (35), schema depth (25), llms.txt (20) and crawler governance (20), capped at 64 when a page arrives as an empty JavaScript shell. Robots rules were evaluated for the home page path.
- Two domains in the original list were corrected (bombas.com, rothys.com). The bare domain redirected our crawler to us.checkout.gymshark.com, so the storefront at www.gymshark.com was audited.
- The full per-domain data, including the validator's reasons for every llms.txt it rejected, is in the CSV and JSON.