---
title: "Which AI crawlers do major sites actually block?"
url: https://proxy.wiki/research/ai-crawler-robots-policies/
type: Research Report
author: "proxy.wiki editorial"
published: 2026-09-01
updated: 2026-09-01
site: proxy.wiki
topics: ["Web scraping"]
license: CC BY 4.0 — quote freely with attribution to https://proxy.wiki/
---

# Which AI crawlers do major sites actually block?

> 46 sites, 18 AI agents, read from robots.txt. ClaudeBot was blocked by 22 and GPTBot by 16, and the ordering is strict.

## Key takeaways

- ClaudeBot was blocked by 22 of 46 sites and GPTBot by 16. Six sites block ClaudeBot without blocking GPTBot; none does the reverse.
- Publishers separate training from answering: GPTBot blocked 16 times against OAI-SearchBot 7, ClaudeBot 22 against Claude-User 12.
- 67 percent of the sample, 31 of 46 sites, name at least one AI agent in robots.txt.
- Only 6 of 46 sites block every crawler through the wildcard group. AI blocks are additions to an otherwise open file.
- Common Crawl's CCBot is blocked as often as the model developers themselves, at 21 sites.
- Four domains refused to serve robots.txt to an ordinary client at all, so their stated policy cannot be read.

The usual assumption is that OpenAI’s crawler is the one publishers block. In this sample it is not. Across 46 well-known sites whose robots.txt we could read, **ClaudeBot was blocked by 22 and GPTBot by 16** — and the difference is not noise, because the relationship is strictly ordered: six sites block ClaudeBot without blocking GPTBot, and not one does the reverse.

This is a snapshot of stated policy, taken on one day. It says nothing about which crawlers obey it.

## Method

We requested `https://<host>/robots.txt` from 50 well-known domains across news, reference, developer, commerce and health, following redirects and identifying ourselves as `proxy.wiki-research/1.0`. 46 returned a readable file.

Each file was parsed into user-agent groups, with consecutive `User-agent:` lines treated as one group sharing the rules that follow, which is what RFC 9309 specifies and what large sites rely on heavily. For each AI agent we recorded one of four verdicts:

- **blocked** — a group naming that agent contains `Disallow: /`

- **allowed** — a group names it and does not disallow everything

- **partial** — specific paths are disallowed, but not the whole site

- **unmentioned** — the file never names that agent

The parser was checked against raw files before the numbers were used. On nytimes.com it reported GPTBot blocked, and the file contains `User-agent: GPTBot` followed by `Disallow: /`. On theguardian.com it reported ClaudeBot blocked from a group of about twenty agents, and that group terminates in `Disallow: /`.

## What the 46 sites say

| Agent | Operator | Blocked | Partial | Allowed | Unmentioned |
| --- | --- | --- | --- | --- | --- |
| ClaudeBot | Anthropic | 22 | 1 | 1 | 22 |
| CCBot | Common Crawl | 21 | 1 | 2 | 22 |
| Applebot-Extended | Apple | 19 | 1 | 1 | 25 |
| Bytespider | ByteDance | 18 | 1 | 1 | 26 |
| meta-externalagent | Meta | 18 | 1 | 0 | 27 |
| Google-Extended | Google | 17 | 1 | 2 | 26 |
| GPTBot | OpenAI | 16 | 3 | 3 | 24 |
| PerplexityBot | Perplexity | 16 | 2 | 2 | 26 |
| Amazonbot | Amazon | 16 | 3 | 1 | 26 |
| anthropic-ai | Anthropic (legacy) | 15 | 1 | 2 | 28 |
| cohere-ai | Cohere | 14 | 0 | 1 | 31 |
| Diffbot | Diffbot | 13 | 1 | 0 | 32 |
| omgili | Webz.io | 13 | 0 | 0 | 33 |
| Claude-User | Anthropic | 12 | 2 | 0 | 32 |
| ChatGPT-User | OpenAI | 11 | 6 | 1 | 28 |
| Timpibot | Timpi | 9 | 1 | 0 | 36 |
| ImagesiftBot | ImageSift | 8 | 0 | 0 | 38 |
| OAI-SearchBot | OpenAI | 7 | 6 | 1 | 32 |

## Finding 1: publishers distinguish training from answering

The clearest pattern in the table is not about which company is least popular. It is that the same operator gets very different treatment depending on what the crawler is _for_.

| Operator | Training or bulk crawl | Fetching to answer a user |
| --- | --- | --- |
| OpenAI | GPTBot — 16 blocked | OAI-SearchBot — 7 blocked |
| Anthropic | ClaudeBot — 22 blocked | Claude-User — 12 blocked |

Both operators see their search or user-triggered agent blocked roughly half as often as their bulk crawler. Publishers are not refusing AI companies outright; they are refusing to be training data while remaining willing to be cited.

The _partial_ column supports this. `OAI-SearchBot` and `ChatGPT-User` each have six sites that restrict some paths rather than the whole site — the highest partial counts in the table. Those sites wrote a considered rule instead of a blanket one.

## Finding 2: ClaudeBot’s block list is a superset of GPTBot’s

Six sites in the sample block ClaudeBot while permitting GPTBot: theguardian.com, washingtonpost.com, arstechnica.com, theverge.com, wired.com and investopedia.com. **No site does the reverse.**

Two things plausibly explain it, and we can only distinguish them with data we do not have. Anthropic publishes several agent names — `ClaudeBot`, `Claude-User`, `Claude-SearchBot`, plus the legacy `anthropic-ai` — and sites that maintain long block lists tend to name all of them, which inflates the count for any operator with more published names. It is also true that several of these files group twenty or more agents under a single `Disallow: /`, so a name being present says more about how the list was assembled than about a decision taken agent by agent.

What the data does support: if you are estimating who can crawl a given publisher, the operator with more published crawler names is the one more likely to be named in a block.

## Finding 3: two thirds of these sites have an AI policy at all

**31 of 46 sites (67%)** name at least one AI agent in robots.txt. Fifteen name none, and are covered only by whatever their wildcard group says.

The wildcard groups themselves are not the story people expect:

| Wildcard User-agent: * posture | Sites |
| --- | --- |
| Partial — some paths disallowed | 37 |
| Blocked — Disallow: / for everyone | 6 |
| No wildcard group at all | 2 |
| Allowed — explicitly everything | 1 |

Only six of 46 close the door to all crawlers. The AI-specific blocks are additions on top of an otherwise open file, aimed at named agents, not a general retreat from being crawled.

## Finding 4: fifteen sites block the big three together

Fifteen sites block `GPTBot`, `ClaudeBot` and `CCBot` simultaneously. That combination — two model developers and the Common Crawl archive that feeds many others — is the closest thing to a standard posture in the sample. Common Crawl being blocked as often as the model developers themselves (21 sites) suggests publishers understand it as an upstream source rather than as an academic archive.

## The 31 sites that name an AI agent

Every site in the sample whose robots.txt names at least one AI crawler, ordered by how many of the eighteen agents it blocks outright. ✗ blocked, ✓ allowed, ~ partial, — unmentioned.

| Site | GPTBot | OAI-Search | ClaudeBot | Google-Ext | CCBot | Perplexity | Bytespider | Blocked of 18 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| nytimes.com | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 |
| cnn.com | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 |
| sciencedirect.com | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | 17 |
| bbc.co.uk | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
| bloomberg.com | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
| amazon.com | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
| theverge.com | ✓ | — | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
| nature.com | ✗ | — | ✗ | ✗ | ✗ | ✗ | ✗ | 15 |
| investopedia.com | ~ | ~ | ✗ | ✗ | ✗ | ✗ | ✗ | 14 |
| arstechnica.com | — | — | ✗ | ✗ | ✗ | ✗ | ✗ | 13 |
| wired.com | — | — | ✗ | ✗ | ✗ | ✗ | ✗ | 12 |
| washingtonpost.com | — | — | ✗ | — | ✗ | ✗ | ✗ | 11 |
| techcrunch.com | ✗ | — | ✗ | ✗ | ✗ | — | ✗ | 11 |
| healthline.com | ✗ | — | ✗ | — | ✗ | — | ✗ | 11 |
| yelp.com | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | — | 10 |
| figma.com | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | — | 10 |
| theguardian.com | — | — | ✗ | — | ✗ | ✗ | ✗ | 9 |
| ebay.com | ✗ | ~ | ✗ | — | ✗ | ✗ | ✗ | 9 |
| tripadvisor.com | ✗ | ~ | ✗ | ✗ | ✗ | ~ | ✗ | 9 |
| medium.com | ✗ | — | ✗ | — | — | — | ✗ | 6 |
| jstor.org | ✗ | — | ✗ | ✗ | ✗ | — | — | 4 |
| webmd.com | ✗ | — | ✗ | — | ✗ | — | — | 3 |
| zillow.com | — | — | — | — | — | — | — | 3 |
| reuters.com | — | ~ | — | — | — | — | — | 1 |
| indeed.com | ~ | ~ | ~ | ~ | ~ | ~ | ~ | 1 |
| notion.so | — | — | — | — | — | — | — | 1 |
| nerdwallet.com | — | — | — | ✗ | — | — | — | 1 |
| wsj.com | ~ | ~ | — | — | — | — | — | 0 |
| imdb.com | — | — | — | — | — | — | — | 0 |
| wordpress.org | ✓ | — | ✓ | ✓ | ✓ | ✓ | ✓ | 0 |
| cloudflare.com | ✓ | — | — | ✓ | ✓ | ✓ | — | 0 |

Three patterns are visible in the rows rather than the totals.

**News sites cluster at the top.** The heaviest blockers are publishers whose product is text: nytimes.com, cnn.com, bbc.co.uk, bloomberg.com. That is the group with both the clearest commercial exposure to a model answering questions from their reporting, and the legal departments to act on it.

**Developer and reference sites cluster at the bottom.** Several name only one or two agents, or none, and are governed entirely by a wildcard group that has not changed in years. Whether that reflects a decision or an absence of one is not something a robots file can tell you.

**The search agents are the exception inside otherwise total blocks.** Look along the OAI-Search column against the GPTBot column: several sites that block the crawler leave the search agent alone. On sciencedirect.com the search agent is explicitly allowed while GPTBot is blocked, which is the training-versus-answering split stated as plainly as a robots file can state it.

## What this does not tell you

robots.txt records a stated wish. It is not enforcement, and RFC 9309 is explicit that compliance is voluntary. A file that names an agent tells you what the operator asked for; it tells you nothing about what that agent did.

Four of the 50 domains — stackoverflow.com, quora.com, pinterest.com and mayoclinic.org — refused our request for the file itself. That is its own finding: on those sites the crawling policy cannot be read by an ordinary client at all, which makes “check robots.txt before you crawl” advice that cannot be followed there.

The counts are also sensitive to how many names an operator publishes, as Finding 2 sets out. A per-operator total would be a different and arguably more honest metric than a per-agent one, and we have not computed it because the mapping from agent to operator is our judgement rather than a published fact.

## If you run a site, what this implies

Three things follow from the data, whichever side of the crawl you are on.

**Naming agents individually is now the norm, and it ages badly.** Thirty-one of these files name specific agents, and the operators keep adding names. A file listing `anthropic-ai` but not `ClaudeBot`, or `GPTBot` but not `OAI-SearchBot`, expresses an intention that the newer agent does not read. Fifteen sites in the sample name `anthropic-ai`, an agent Anthropic has superseded, which is a reasonable proxy for how many of these lists are maintained by hand and revisited rarely.

**The training-versus-answering distinction is the decision worth making deliberately.** Blocking a bulk crawler keeps your text out of a training set. Blocking the agent that fetches a page to answer a live question removes you from the answer instead, along with any citation it would have carried. Those are opposite outcomes, and a single `Disallow: /` across every AI agent chooses both without distinguishing them. The sites in this sample that wrote partial rules were, almost without exception, drawing exactly that line.

**A file nobody can fetch states nothing.** Four of the fifty domains refused our request. If your edge rules block unfamiliar clients before they reach `/robots.txt`, a compliant crawler cannot discover your policy and a non-compliant one is unaffected — the check falls only on the crawlers that were going to obey it.

Our own file takes the opposite position and says so in a comment: full-text crawling is welcome, including for training and answer engines. That is a choice a reference site can make and a newspaper cannot, which is most of what the table above shows.

## Related reading

The mechanics of the file itself, including the longest-match rule this study does not evaluate, are in the [robots.txt glossary entry](/glossary/robots-txt/). What the file is and is not able to enforce sits alongside the wider legal picture in [web scraping](/glossary/web-scraping/).

For the layers below stated policy — what a server can tell about a client before it reads a single header — see our companion measurements on [what HTTP clients send by default](/research/default-request-headers/) and on [TLS handshake fingerprints](/research/tls-clienthello-fingerprints/). Between them: robots.txt records what a site asks, and those two record what it can detect regardless.

## Reproducing this

The whole study is a fetch and a parse. The parsing rule that matters is the grouping one — consecutive `User-agent:` lines share the rules that follow them, and a parser that misses this will under-count every site that maintains a long list.

```
def groups(txt):
    """user-agent (lower) -> [(directive, path)]"""
    out, current, prev_ua = {}, [], False
    for raw in txt.splitlines():
        line = raw.split('#', 1)[0].strip()
        if not line or ':' not in line:
            continue
        k, v = [x.strip() for x in line.split(':', 1)]
        k = k.lower()
        if k == 'user-agent':
            if not prev_ua:          # a new block starts here
                current = []
            current.append(v.lower())
            for ua in current:
                out.setdefault(ua, [])
            prev_ua = True
        elif k in ('allow', 'disallow'):
            prev_ua = False
            for ua in current:
                out.setdefault(ua, []).append((k, v))
    return out
```

Verify any parser against a raw file before trusting its totals. Ours was checked against two sites with different structures — one naming a single agent, one grouping twenty — and both agreed.

## Limits

- **One day, 50 domains.** Fetched on 2 September 2026. These files change often; several carry comments describing recent policy changes.

- **A hand-picked sample, not a ranking.** The 50 were chosen to span news, reference, developer, commerce and health. They are not the top 50 by traffic and the percentages should not be read as representative of the web.

- **Site-wide posture only.** We record whether an agent is disallowed from everything, not which paths a partial rule covers.

- **Stated policy, not behaviour.** Nothing here measures whether any crawler complied.

## Frequently asked questions

### Which AI crawler is blocked most often?

In this sample of 46 sites, ClaudeBot, at 22 sites, ahead of Common Crawl's CCBot at 21 and GPTBot at 16. The ordering is consistent rather than marginal: six sites block ClaudeBot while permitting GPTBot, and no site does the opposite.

### Does that mean publishers object to Anthropic more than to OpenAI?

Not necessarily, and the article says so. Anthropic publishes more crawler names than OpenAI, and sites maintaining long block lists tend to name all of an operator's agents. A per-operator count would be a fairer measure than a per-agent one, and it depends on a mapping we would be asserting rather than citing.

### Do sites block AI crawlers and search crawlers equally?

No, and this is the clearest pattern in the data. Agents that fetch a page to answer a user's question are blocked roughly half as often as agents that crawl in bulk for training. OpenAI's OAI-SearchBot was blocked by 7 sites against GPTBot's 16; Anthropic's Claude-User by 12 against ClaudeBot's 22.

### Does robots.txt actually stop a crawler?

No. It is a voluntary protocol, standardised in RFC 9309, that records what a site operator asks crawlers to do. This study measures stated policy only and makes no claim about whether any crawler complied.

### How many sites have no AI policy at all?

Fifteen of the 46 name no AI agent, so they are governed only by their wildcard group. Of the whole sample, just six block every crawler through that wildcard, meaning most files remain open to general crawling and add named exceptions on top.

### What is the most common parsing mistake with robots.txt?

Treating each User-agent line as starting a new group. Consecutive User-agent lines share the rules that follow them, and large sites rely on this heavily. One site in this sample groups about twenty agents under a single Disallow, and a parser that missed the grouping rule would score all but the last of them as unmentioned.

## Sources

1. [RFC 9309: Robots Exclusion Protocol, including the grouping rules](https://www.rfc-editor.org/rfc/rfc9309.html)
2. [OpenAI: the crawler names it publishes and what each is for](https://platform.openai.com/docs/bots)
3. [Dark Visitors: a public directory of AI agent user agents](https://darkvisitors.com/agents)
