Collecting text at the scale a language model needs is not a harder version of scraping one site. It is a different problem, and the parts that break are not the parts people prepare for. The fetching is the easy half. Deduplication, provenance and knowing when to stop are what decide whether the corpus is usable.
This guide covers the operational side: what to record per document, how to respect a policy you can actually read, and the failure modes that only appear past a few million pages.
#Read the policy before you read the page
robots.txt is the only machine-readable statement a site makes about crawling, and since 2022 it has been a standard rather than a convention. It is advisory, not enforcement — but for a training corpus the question is not whether you can be stopped. It is whether you can show what you were told.
Two details of the format cause most mistakes, and both are in RFC 9309.
Consecutive user-agent lines share the rules that follow them. A parser that treats each User-agent: line as starting a new group will misread most large sites. We measured this: in a survey of 46 major sites, one groups about twenty agents under a single Disallow: /. A naive parser scores nineteen of them as unmentioned.
The most specific matching group wins, and only that group applies. If a file names your agent, the wildcard group is irrelevant to you — including any Allow lines in it.
# Your agent matches this group only. The * group below does not apply.
User-agent: MyCrawler
Disallow: /
User-agent: *
Allow: /
That survey also found something worth knowing before you choose a user-agent string: publishers block bulk crawlers roughly twice as often as they block agents that fetch a page to answer a live question. Sixteen of 46 sites blocked GPTBot; seven blocked OAI-SearchBot. If your crawler is genuinely doing one and not the other, saying so in the name is not cosmetic.
#Record provenance at fetch time or lose it
The single most expensive mistake is treating the corpus as a pile of text. Every downstream question — can we use this, where did it come from, is it stale, was it duplicated — needs metadata that is cheap to record at fetch time and impossible to reconstruct later.
| Field | Why it is needed later |
|---|---|
| URL, final after redirects | The URL you requested is often not the one you got |
| Fetch timestamp, UTC | Freshness, and evidence of what the policy was at the time |
| HTTP status and content type | A 200 serving HTML when you expected JSON is a silent corruption |
| robots.txt verdict at fetch time | Policies change; what you were told then is not recoverable now |
| Content hash | Deduplication, and detecting silent edits on re-crawl |
| Response byte length | Catches truncation that parsing will not |
| Canonical URL from the page | The site’s own statement about which copy is authoritative |
Store these beside the document, not in a separate log that will be rotated away. A corpus without per-document provenance cannot be filtered, audited or partially deleted, and all three get asked for eventually.
#Deduplication is most of the work
The web is enormously repetitive, and a training corpus that does not deduplicate is mostly boilerplate. Three layers, cheapest first:
- Exact hash. A hash of the normalised body catches byte-identical duplicates: mirrors, syndication, the same page under two URLs. Cheap and catches a great deal.
- URL canonicalisation. Strip tracking parameters, respect the page’s own
rel="canonical", and normalise trailing slashes and case in the host. Much duplication is one document behind many addresses. - Near-duplicate detection. Templated pages differ only in a product name or a date. This is the expensive layer and the one that decides quality; techniques such as MinHash or SimHash exist precisely because exact hashing misses it.
Do the first two at write time. They cost almost nothing and shrink what the third layer has to examine.
#Rate, and why the address is not the fix
The instinct at scale is to add addresses. It works for per-address rate limits and for nothing else. Our measurements of what HTTP clients send by default and how their TLS handshakes differ both point the same way: most of what identifies a bulk client is produced by the client, not the network path, and a new address changes none of it.
What does help:
# One host at a time, with a delay proportional to its response time.
# A slow server is telling you something; a fixed delay ignores it.
delay = max(min_delay, last_response_seconds * courtesy_factor)
# Honour Retry-After exactly. It is standardised and it is an instruction.
if response.status == 429 and 'Retry-After' in response.headers:
sleep(parse_retry_after(response.headers['Retry-After']))
Per-host concurrency of one or two, with backoff keyed to the host’s own latency, will collect more over a week than aggressive parallelism collects before it gets refused. Concurrency is the setting that causes most blocks, not total volume.
#Identify yourself, and mean it
A crawler’s user agent is a claim, and at this scale it is a claim you will be held to. Three positions are defensible; a fourth is not.
| Position | What it means in practice |
|---|---|
| A named agent with a contact URL | Operators can reach you, and can allow you specifically. The only position that lets a site say yes |
| An honest library default | Unremarkable and truthful. Fine for small, permitted work |
| No user agent at all | Treated as more suspicious than an unrecognised one, because ordinary traffic almost always carries something |
| A copied browser string | Contradicted by everything else your client does, and it removes any claim that access was in good faith |
User-Agent: ExampleCorpBot/1.0 (+https://example.com/bot)
The fourth position is worth spelling out because it is the common one. Our measurements show a browser sending eleven headers where the median library sends four, in a different order, with none of the Sec-Fetch-* set — so a browser string over a library connection is not a disguise. It is a contradiction, and a contradiction is more distinctive than an honest default. The same holds one layer down: the TLS handshake is fixed per library and cannot be changed by anything at the HTTP layer.
If you want the option of being permitted rather than merely undetected, a named agent is the only route. It is also the only one that survives someone asking, later, what you were doing.
#Storage shape decides what you can do later
Two decisions taken on day one are expensive to reverse.
Store the raw response, not just the extracted text. Extraction rules improve, and re-extracting from stored HTML costs nothing while re-fetching a million pages costs a great deal and may no longer be possible. Compressed raw HTML is not the expensive part of a corpus.
Partition by source and by fetch date. Every operational request you will receive later — remove this domain, rebuild everything older than a date, audit what came from one site — is trivial against that layout and painful against a flat store. It costs nothing at write time.
Both follow the same principle as the provenance fields above: record what is cheap now and irrecoverable later.
#What breaks past a few million pages
These are the failures that do not appear in a pilot run.
| Failure | Symptom | Guard |
|---|---|---|
| Silent content drift | A parser change quietly empties a field for one site | Per-source populated-field counts, compared run to run |
| Challenge pages stored as text | Thousands of near-identical “verify you are human” documents | Assert on content, not on a 200 status |
| Redirect loops into one page | One document, thousands of URLs | Hash before store; record the final URL |
| Encoding damage | Mojibake in a fraction of documents | Trust the declared charset, verify by decode |
| Unbounded queue growth | Memory exhaustion days in | Bound the frontier; persist it |
| Politeness state lost on restart | A restart hammers hosts you had backed off from | Persist per-host backoff, not just the queue |
The first two are the expensive ones because they are invisible. A pipeline storing challenge pages reports a healthy success rate the whole time. This is why success measured on status codes is misleading: a 200 asserts that a response was returned, not that it was the document you asked for.
#Decide what you will not collect
Some of this is legal and jurisdiction-specific, and none of it is advice. But three categories are worth an explicit decision before the crawl starts rather than after:
- Personal data. Obligations attach to it regardless of whether the page was public. Being able to find and delete every document from one source later is a design requirement, not a feature.
- Material behind a login or a paywall. Public and accessible are different things, and the distinction matters more here than almost anywhere.
- Sites that asked you not to. Whether or not the file binds you, ignoring it after reading it is a different position from never having looked.
The provenance fields above are what make any of this actionable. Without a per-document record of source and fetch time, “remove everything from that domain” is not a query you can run.
#A note on being crawled yourself
If you publish as well as collect, the same survey is worth reading from the other side. Two thirds of the sites we checked name at least one AI agent in robots.txt, and the ones writing careful rules are drawing the training-versus-answering line deliberately rather than blocking everything. Whichever side you are on, that distinction is the one worth being precise about.
#Where to go next
The mechanics of the exclusion protocol are in the robots.txt entry. Rate limiting, backoff and what Retry-After obliges you to do are covered in reading proxy and block errors. For how a bulk client is identified before it sends a header, see the two measurement studies linked above.