The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSBENCHMARKSDATALEARNNEWSLETTERPARTNER WITH US →
Data/Web Ecosystem

Web Ecosystem

AI Crawler Blocking in Top Websites

Share of observed sites in a frozen current Tranco top-1,000 cohort whose robots.txt declares a full block for six AI crawlers. Applying today’s cohort backward creates selection bias: some members were not prominent in 2023 and past leaders that fell out are absent. It measures policy, not compliance or enforcement.

AI Crawler Blocking in Top Websites

  • GPTBot
  • ClaudeBot
  • CCBot
  • PerplexityBot
  • Google-Extended
  • Bytespider
  • Observed cohort coverage
For 2026-09-07, Observed cohort coverage was blocked by 52.7% of sites with an observed robots.txt; consult the coverage line before comparing periods.60%40%20%0%Aug '26Sep '26{"f":[800,420,56,16],"s":[["GPTBot","#10A37F"],["ClaudeBot","#D97757"],["CCBot","#7C3AED"],["PerplexityBot","#0EA5A4"],["Google-Extended","#4285F4"],["Bytespider","#E45756"],["Observed cohort coverage","#64748B"]],"p":[["2026-08-01T00:00:00.000Z","Aug '26",56,[[6,"13.4%",298.71,null]],null],["2026-08-27T00:00:00.000Z","Aug '26",567.57,[[0,"21.68%",248.46,null],[1,"22.43%",243.93,null],[2,"25.23%",226.92,null],[3,"17.76%",272.27,null],[4,"20%",258.67,null],[5,"24.3%",232.59,null],[6,"53.5%",55.43,null]],null],["2026-09-07T00:00:00.000Z","Sep '26",784,[[0,"21.63%",248.77,null],[1,"21.63%",248.77,null],[2,"24.67%",230.35,null],[3,"17.84%",271.79,null],[4,"19.73%",260.28,null],[5,"23.91%",234.95,null],[6,"52.7%",60.29,null]],null]]}GPTBot: 21.6% on Sep '26. ClaudeBot: 21.6% on Sep '26. CCBot: 24.7% on Sep '26. PerplexityBot: 17.8% on Sep '26. Google-Extended: 19.7% on Sep '26. Bytespider: 23.9% on Sep '26. Observed cohort coverage: 52.7% on Sep '26
SOURCE: Internet Archive Wayback CDX and fixed Tranco 46W9X cohort
RANGEALLYTD12M3M1M
Key takeaway

For 2026-09-07, Observed cohort coverage was blocked by 52.7% of sites with an observed robots.txt; consult the coverage line before comparing periods.

For 2026-09-07, Observed cohort coverage was blocked by 52.7% of sites with an observed robots.txt; consult the coverage line before comparing periods.

Pro API coming soon

Methodology

The cohort is the first 1,000 pay-level domains in Tranco list 46W9X, generated on August 26, 2026 from provider ranks spanning July 28–August 26 and retrieved on 2026-08-27. The exact permanent list is https://tranco-list.eu/list/46W9X; the committed cohort artifact has SHA-256 3f9757baf8b2d6a9cc305a01b9c4be4074b8da39ffd54b7a473badc38b17adc1. Membership and rank remain fixed for every historical and forward observation. Applying this current cohort backward creates survivorship and selection bias: some current top-1,000 domains were not prominent in 2023, while historical leaders that later fell out of the ranking are absent. It is not the contemporaneous top 1,000 in each period.

Before the full backfill, the updater queried the first 20 cohort domains monthly from August 2023 through 2026-08-27. For each domain it made two Internet Archive Wayback CDX requests spanning the pilot window with collapse=timestamp:6: the forward result supplied the first HTTP-200 capture in each month and the reverse result supplied the last. For each first-of-month target, those candidates contain the earliest capture after the target and latest capture before it. The updater selected the closer candidate within ±14 days, using CDX URL canonicalization for HTTP/HTTPS and www/apex variants, fetched the raw id_ replay, and counted a site as observed only when that body was retrievable and parseable as robots.txt. Exact pilot period coverage is committed in robots-txt-ai-crawler-blocking-rate-pilot-coverage.csv.

Pilot coverage was partial—generally 55–70% through 2025 and lower in recent months—so the full historical cohort uses quarterly targets rather than presenting a denser monthly line with the same nonrandom archive gaps. The targets are August 2023, November 2023, and three-month intervals through August 2026. Each timestamp is the target date, not the capture timestamp. Only captures within ±14 days qualify. Full monthly retrieval would also require validating as many as 37,000 replay bodies, while quarterly retrieval requires 13,000.

The denominator for each bot share is only cohort sites with a successfully fetched, usable robots.txt snapshot for that period. A missing CDX record, CDX timeout, failed replay, HTML error body, non-200 response, or failed live request is excluded from both numerator and denominator; it is never coded as allowing a crawler. Observed cohort coverage is published as its own series for every target and is also stated in every point note. Archived coverage ranged from 10.8% (2026-05-01) to 50.5%. Bot-rate points are suppressed when coverage is below 30%; this removes May and August 2026 rate estimates affected by recent-capture lag while preserving their coverage evidence.

The parser follows RFC 9309 group selection and rule precedence for the crawler product tokens GPTBot, ClaudeBot, CCBot, PerplexityBot, Google-Extended, and Bytespider. Matching is case-insensitive. The most specific matching user-agent group is used; wildcard User-agent: * applies only when no more-specific group matches; repeated equally specific groups are combined; empty Disallow values do not block; and the longest matching Allow/Disallow pattern wins, with Allow winning equal-length ties. Wildcard * and terminal $ path patterns are supported.

A bot is classified as blocked only when the applicable rules disallow the root path “/”, meaning a full-site robots exclusion. Partial exclusions such as Disallow: /search are not counted as a full block. Robots.txt measures declared machine-readable policy, not crawler compliance or enforcement. WAF rules, authentication, rate limits, and network blocks can restrict access without a robots.txt rule, so the metric can understate total access controls; conversely, declared rules may not be enforced or honored, so the rate is not a strict numeric floor on effective denial. Different crawler operators may interpret nonstandard syntax differently.

The forward collector requests https://DOMAIN/robots.txt for the same fixed cohort each week, follows redirects, and falls back to HTTP only when HTTPS does not produce a usable HTTP-200 robots body. It appends or replaces that collection date while preserving the quarterly archive. Live collection and Wayback history have different availability mechanisms, so the coverage line and the August 2026 transition should be considered when comparing them.

Frequently asked questions

Which sites are in the cohort, and why is it fixed?

The cohort is exactly the first 1,000 pay-level domains in permanent Tranco list 46W9X, generated on August 26, 2026 and retrieved on August 27. Membership and rank stay fixed across every period. Applying this current cohort backward creates survivorship and selection bias: some members were not prominent in 2023, while historical leaders that later fell out are absent. The chart therefore does not represent each period’s contemporaneous top 1,000.

Why do AI crawler blocking rates matter?

Robots.txt is a widely used way for publishers to declare whether automated agents may crawl their sites. Changes in full-site blocking indicate how publishers are trying to control access that may support AI training or search and answer products, illuminating negotiations and friction between content owners, AI companies, and web platforms. The rates do not measure actual crawling or content use.

Which six crawler user agents are measured?

The parser tests the product tokens GPTBot, ClaudeBot, CCBot, PerplexityBot, Google-Extended, and Bytespider. They are shown separately because operators publish different crawler identities and may use them for different AI training, indexing, search, or product functions; inclusion does not imply that all six have identical purposes or behavior.

How does the parser interpret groups, wildcards, and conflicting rules?

Matching is case-insensitive and follows RFC 9309: the most specific matching user-agent group applies, equally specific repeated groups are combined, and User-agent: * applies only when no more-specific group matches. Empty Disallow values do not block. The longest matching Allow or Disallow path wins, with Allow winning equal-length ties; * wildcards and terminal $ anchors are supported. A site counts as fully blocked only when the applicable result denies the root path /.

How are historical Wayback snapshots selected, and why quarterly?

For each quarterly target, the collector compares the nearest qualifying HTTP-200 robots.txt captures before and after the target and uses the closer usable snapshot within ±14 days. The timestamp shown is the target date, not the capture time. A 20-domain monthly pilot found partial, nonrandom archive coverage, so quarterly points avoid implying monthly precision and reduce full-cohort replay validation from as many as 37,000 to 13,000 bodies.

What is the denominator, and what does the coverage series show?

Each bot’s rate uses only cohort sites with a successfully fetched, usable robots.txt observation for that period. Missing captures, timeouts, failed replays, HTML error bodies, non-200 responses, and failed live requests are excluded from both numerator and denominator—not counted as allowing a crawler. The separate coverage series reports observed sites as a share of all 1,000 cohort domains, and point notes give the exact count.

Why are some low-coverage bot-rate points suppressed?

Bot rates are not published when fewer than 30% of cohort sites have usable observations. Wayback coverage was 10.8% for May 2026 and 13.4% for August 2026, so those bot-rate points are suppressed because the observed subset may be especially unrepresentative. Their coverage points remain visible as evidence of the archive gap.

Why does the series begin in August 2023?

GPTBot’s introduction in August 2023 provides a natural starting month for tracking explicit AI-crawler policy. The first target is August 1 and can use a qualifying capture within ±14 days, so it should be read as an August snapshot rather than a precise launch-day measurement. No earlier values are backfilled as zeros. Rules matching User-agent: * may also produce a parsed result even before a named token was widely used.

Does robots.txt guarantee enforcement or establish a legal prohibition?

No. Robots.txt records declared machine-readable policy; it does not enforce access or show crawler compliance. WAF rules, authentication, rate limits, and network blocks can deny access without a robots.txt rule, so this metric can understate total access controls. Conversely, declared rules may not be enforced or honored, so the rate is not a strict numeric floor on effective denial. It also does not establish legal permission or prohibition.

How will the chart be updated going forward?

The forward collector requests each frozen-cohort domain’s live robots.txt weekly, following redirects and falling back to HTTP only when HTTPS does not return a usable HTTP-200 robots body. It appends or replaces that collection date without changing cohort membership or the quarterly archive. Because live and Wayback availability differ, compare the transition using the coverage series.

How should rates be compared across crawler operators?

Treat each line as the share of observed sites whose text parses as a full-root block for that named token, not as a ranking of operator compliance or access. Publishers may target operators differently, wildcard rules can affect several tokens at once, crawler purposes differ, and operators may interpret nonstandard syntax differently. Compare direction and gaps alongside coverage, launch timing, and operator documentation.

Related charts

  • AI Crawler Traffic Share by Operator

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • LEARN
  • Data
  • Benchmarks
  • Score
  • Methodology
  • About
  • Team
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block
  • Add The Latent as a preferred source