The usual assumption is that OpenAI’s crawler is the one publishers block. In this sample it is not. Across 46 well-known sites whose robots.txt we could read, ClaudeBot was blocked by 22 and GPTBot by 16 — and the difference is not noise, because the relationship is strictly ordered: six sites block ClaudeBot without blocking GPTBot, and not one does the reverse.
This is a snapshot of stated policy, taken on one day. It says nothing about which crawlers obey it.
#Method
We requested https://<host>/robots.txt from 50 well-known domains across news, reference, developer, commerce and health, following redirects and identifying ourselves as proxy.wiki-research/1.0. 46 returned a readable file.
Each file was parsed into user-agent groups, with consecutive User-agent: lines treated as one group sharing the rules that follow, which is what RFC 9309 specifies and what large sites rely on heavily. For each AI agent we recorded one of four verdicts:
- blocked — a group naming that agent contains
Disallow: / - allowed — a group names it and does not disallow everything
- partial — specific paths are disallowed, but not the whole site
- unmentioned — the file never names that agent
The parser was checked against raw files before the numbers were used. On nytimes.com it reported GPTBot blocked, and the file contains User-agent: GPTBot followed by Disallow: /. On theguardian.com it reported ClaudeBot blocked from a group of about twenty agents, and that group terminates in Disallow: /.
#What the 46 sites say
| Agent | Operator | Blocked | Partial | Allowed | Unmentioned |
|---|---|---|---|---|---|
| ClaudeBot | Anthropic | 22 | 1 | 1 | 22 |
| CCBot | Common Crawl | 21 | 1 | 2 | 22 |
| Applebot-Extended | Apple | 19 | 1 | 1 | 25 |
| Bytespider | ByteDance | 18 | 1 | 1 | 26 |
| meta-externalagent | Meta | 18 | 1 | 0 | 27 |
| Google-Extended | 17 | 1 | 2 | 26 | |
| GPTBot | OpenAI | 16 | 3 | 3 | 24 |
| PerplexityBot | Perplexity | 16 | 2 | 2 | 26 |
| Amazonbot | Amazon | 16 | 3 | 1 | 26 |
| anthropic-ai | Anthropic (legacy) | 15 | 1 | 2 | 28 |
| cohere-ai | Cohere | 14 | 0 | 1 | 31 |
| Diffbot | Diffbot | 13 | 1 | 0 | 32 |
| omgili | Webz.io | 13 | 0 | 0 | 33 |
| Claude-User | Anthropic | 12 | 2 | 0 | 32 |
| ChatGPT-User | OpenAI | 11 | 6 | 1 | 28 |
| Timpibot | Timpi | 9 | 1 | 0 | 36 |
| ImagesiftBot | ImageSift | 8 | 0 | 0 | 38 |
| OAI-SearchBot | OpenAI | 7 | 6 | 1 | 32 |
#Finding 1: publishers distinguish training from answering
The clearest pattern in the table is not about which company is least popular. It is that the same operator gets very different treatment depending on what the crawler is for.
| Operator | Training or bulk crawl | Fetching to answer a user |
|---|---|---|
| OpenAI | GPTBot — 16 blocked | OAI-SearchBot — 7 blocked |
| Anthropic | ClaudeBot — 22 blocked | Claude-User — 12 blocked |
Both operators see their search or user-triggered agent blocked roughly half as often as their bulk crawler. Publishers are not refusing AI companies outright; they are refusing to be training data while remaining willing to be cited.
The partial column supports this. OAI-SearchBot and ChatGPT-User each have six sites that restrict some paths rather than the whole site — the highest partial counts in the table. Those sites wrote a considered rule instead of a blanket one.
#Finding 2: ClaudeBot’s block list is a superset of GPTBot’s
Six sites in the sample block ClaudeBot while permitting GPTBot: theguardian.com, washingtonpost.com, arstechnica.com, theverge.com, wired.com and investopedia.com. No site does the reverse.
Two things plausibly explain it, and we can only distinguish them with data we do not have. Anthropic publishes several agent names — ClaudeBot, Claude-User, Claude-SearchBot, plus the legacy anthropic-ai — and sites that maintain long block lists tend to name all of them, which inflates the count for any operator with more published names. It is also true that several of these files group twenty or more agents under a single Disallow: /, so a name being present says more about how the list was assembled than about a decision taken agent by agent.
What the data does support: if you are estimating who can crawl a given publisher, the operator with more published crawler names is the one more likely to be named in a block.
#Finding 3: two thirds of these sites have an AI policy at all
31 of 46 sites (67%) name at least one AI agent in robots.txt. Fifteen name none, and are covered only by whatever their wildcard group says.
The wildcard groups themselves are not the story people expect:
Wildcard User-agent: * posture |
Sites |
|---|---|
| Partial — some paths disallowed | 37 |
Blocked — Disallow: / for everyone |
6 |
| No wildcard group at all | 2 |
| Allowed — explicitly everything | 1 |
Only six of 46 close the door to all crawlers. The AI-specific blocks are additions on top of an otherwise open file, aimed at named agents, not a general retreat from being crawled.
#Finding 4: fifteen sites block the big three together
Fifteen sites block GPTBot, ClaudeBot and CCBot simultaneously. That combination — two model developers and the Common Crawl archive that feeds many others — is the closest thing to a standard posture in the sample. Common Crawl being blocked as often as the model developers themselves (21 sites) suggests publishers understand it as an upstream source rather than as an academic archive.
#The 31 sites that name an AI agent
Every site in the sample whose robots.txt names at least one AI crawler, ordered by how many of the eighteen agents it blocks outright. ✗ blocked, ✓ allowed, ~ partial, — unmentioned.
| Site | GPTBot | OAI-Search | ClaudeBot | Google-Ext | CCBot | Perplexity | Bytespider | Blocked of 18 |
|---|---|---|---|---|---|---|---|---|
nytimes.com |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 |
cnn.com |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 |
sciencedirect.com |
✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 |
bbc.co.uk |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
bloomberg.com |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
amazon.com |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
theverge.com |
✓ | — | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
nature.com |
✗ | — | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
investopedia.com |
~ | ~ | ✗ | ✗ | ✗ | ✗ | ✗ | 14 |
arstechnica.com |
— | — | ✗ | ✗ | ✗ | ✗ | ✗ | 13 |
wired.com |
— | — | ✗ | ✗ | ✗ | ✗ | ✗ | 12 |
washingtonpost.com |
— | — | ✗ | — | ✗ | ✗ | ✗ | 11 |
techcrunch.com |
✗ | — | ✗ | ✗ | ✗ | — | ✗ | 11 |
healthline.com |
✗ | — | ✗ | — | ✗ | — | ✗ | 11 |
yelp.com |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | — | 10 |
figma.com |
✗ | ✗ | ✗ | ✗ | ✗ | ✗ | — | 10 |
theguardian.com |
— | — | ✗ | — | ✗ | ✗ | ✗ | 9 |
ebay.com |
✗ | ~ | ✗ | — | ✗ | ✗ | ✗ | 9 |
tripadvisor.com |
✗ | ~ | ✗ | ✗ | ✗ | ~ | ✗ | 9 |
medium.com |
✗ | — | ✗ | — | — | — | ✗ | 6 |
jstor.org |
✗ | — | ✗ | ✗ | ✗ | — | — | 4 |
webmd.com |
✗ | — | ✗ | — | ✗ | — | — | 3 |
zillow.com |
— | — | — | — | — | — | — | 3 |
reuters.com |
— | ~ | — | — | — | — | — | 1 |
indeed.com |
~ | ~ | ~ | ~ | ~ | ~ | ~ | 1 |
notion.so |
— | — | — | — | — | — | — | 1 |
nerdwallet.com |
— | — | — | ✗ | — | — | — | 1 |
wsj.com |
~ | ~ | — | — | — | — | — | 0 |
imdb.com |
— | — | — | — | — | — | — | 0 |
wordpress.org |
✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | 0 |
cloudflare.com |
✓ | — | — | ✓ | ✓ | ✓ | — | 0 |
Three patterns are visible in the rows rather than the totals.
News sites cluster at the top. The heaviest blockers are publishers whose product is text: nytimes.com, cnn.com, bbc.co.uk, bloomberg.com. That is the group with both the clearest commercial exposure to a model answering questions from their reporting, and the legal departments to act on it.
Developer and reference sites cluster at the bottom. Several name only one or two agents, or none, and are governed entirely by a wildcard group that has not changed in years. Whether that reflects a decision or an absence of one is not something a robots file can tell you.
The search agents are the exception inside otherwise total blocks. Look along the OAI-Search column against the GPTBot column: several sites that block the crawler leave the search agent alone. On sciencedirect.com the search agent is explicitly allowed while GPTBot is blocked, which is the training-versus-answering split stated as plainly as a robots file can state it.
#What this does not tell you
robots.txt records a stated wish. It is not enforcement, and RFC 9309 is explicit that compliance is voluntary. A file that names an agent tells you what the operator asked for; it tells you nothing about what that agent did.
Four of the 50 domains — stackoverflow.com, quora.com, pinterest.com and mayoclinic.org — refused our request for the file itself. That is its own finding: on those sites the crawling policy cannot be read by an ordinary client at all, which makes “check robots.txt before you crawl” advice that cannot be followed there.
The counts are also sensitive to how many names an operator publishes, as Finding 2 sets out. A per-operator total would be a different and arguably more honest metric than a per-agent one, and we have not computed it because the mapping from agent to operator is our judgement rather than a published fact.
#If you run a site, what this implies
Three things follow from the data, whichever side of the crawl you are on.
Naming agents individually is now the norm, and it ages badly. Thirty-one of these files name specific agents, and the operators keep adding names. A file listing anthropic-ai but not ClaudeBot, or GPTBot but not OAI-SearchBot, expresses an intention that the newer agent does not read. Fifteen sites in the sample name anthropic-ai, an agent Anthropic has superseded, which is a reasonable proxy for how many of these lists are maintained by hand and revisited rarely.
The training-versus-answering distinction is the decision worth making deliberately. Blocking a bulk crawler keeps your text out of a training set. Blocking the agent that fetches a page to answer a live question removes you from the answer instead, along with any citation it would have carried. Those are opposite outcomes, and a single Disallow: / across every AI agent chooses both without distinguishing them. The sites in this sample that wrote partial rules were, almost without exception, drawing exactly that line.
A file nobody can fetch states nothing. Four of the fifty domains refused our request. If your edge rules block unfamiliar clients before they reach /robots.txt, a compliant crawler cannot discover your policy and a non-compliant one is unaffected — the check falls only on the crawlers that were going to obey it.
Our own file takes the opposite position and says so in a comment: full-text crawling is welcome, including for training and answer engines. That is a choice a reference site can make and a newspaper cannot, which is most of what the table above shows.
#Related reading
The mechanics of the file itself, including the longest-match rule this study does not evaluate, are in the robots.txt glossary entry. What the file is and is not able to enforce sits alongside the wider legal picture in web scraping.
For the layers below stated policy — what a server can tell about a client before it reads a single header — see our companion measurements on what HTTP clients send by default and on TLS handshake fingerprints. Between them: robots.txt records what a site asks, and those two record what it can detect regardless.
#Reproducing this
The whole study is a fetch and a parse. The parsing rule that matters is the grouping one — consecutive User-agent: lines share the rules that follow them, and a parser that misses this will under-count every site that maintains a long list.
def groups(txt):
"""user-agent (lower) -> [(directive, path)]"""
out, current, prev_ua = {}, [], False
for raw in txt.splitlines():
line = raw.split('#', 1)[0].strip()
if not line or ':' not in line:
continue
k, v = [x.strip() for x in line.split(':', 1)]
k = k.lower()
if k == 'user-agent':
if not prev_ua: # a new block starts here
current = []
current.append(v.lower())
for ua in current:
out.setdefault(ua, [])
prev_ua = True
elif k in ('allow', 'disallow'):
prev_ua = False
for ua in current:
out.setdefault(ua, []).append((k, v))
return out
Verify any parser against a raw file before trusting its totals. Ours was checked against two sites with different structures — one naming a single agent, one grouping twenty — and both agreed.
#Limits
- One day, 50 domains. Fetched on 2 September 2026. These files change often; several carry comments describing recent policy changes.
- A hand-picked sample, not a ranking. The 50 were chosen to span news, reference, developer, commerce and health. They are not the top 50 by traffic and the percentages should not be read as representative of the web.
- Site-wide posture only. We record whether an agent is disallowed from everything, not which paths a partial rule covers.
- Stated policy, not behaviour. Nothing here measures whether any crawler complied.