Skip to content
AI and agents

Collecting training data for a language model

Provenance, deduplication, rate limits, and the 200 that is not the document — what a bulk text pipeline must record at fetch time or lose for good.

Collecting text at the scale a language model needs is not a harder version of scraping one site. It is a different problem, and the parts that break are not the parts people prepare for. The fetching is the easy half. Deduplication, provenance and knowing when to stop are what decide whether the corpus is usable.

This guide covers the operational side: what to record per document, how to respect a policy you can actually read, and the failure modes that only appear past a few million pages.

#Read the policy before you read the page

robots.txt is the only machine-readable statement a site makes about crawling, and since 2022 it has been a standard rather than a convention. It is advisory, not enforcement — but for a training corpus the question is not whether you can be stopped. It is whether you can show what you were told.

Two details of the format cause most mistakes, and both are in RFC 9309.

Consecutive user-agent lines share the rules that follow them. A parser that treats each User-agent: line as starting a new group will misread most large sites. We measured this: in a survey of 46 major sites, one groups about twenty agents under a single Disallow: /. A naive parser scores nineteen of them as unmentioned.

The most specific matching group wins, and only that group applies. If a file names your agent, the wildcard group is irrelevant to you — including any Allow lines in it.

# Your agent matches this group only. The * group below does not apply.
User-agent: MyCrawler
Disallow: /

User-agent: *
Allow: /

That survey also found something worth knowing before you choose a user-agent string: publishers block bulk crawlers roughly twice as often as they block agents that fetch a page to answer a live question. Sixteen of 46 sites blocked GPTBot; seven blocked OAI-SearchBot. If your crawler is genuinely doing one and not the other, saying so in the name is not cosmetic.

#Record provenance at fetch time or lose it

The single most expensive mistake is treating the corpus as a pile of text. Every downstream question — can we use this, where did it come from, is it stale, was it duplicated — needs metadata that is cheap to record at fetch time and impossible to reconstruct later.

Field Why it is needed later
URL, final after redirects The URL you requested is often not the one you got
Fetch timestamp, UTC Freshness, and evidence of what the policy was at the time
HTTP status and content type A 200 serving HTML when you expected JSON is a silent corruption
robots.txt verdict at fetch time Policies change; what you were told then is not recoverable now
Content hash Deduplication, and detecting silent edits on re-crawl
Response byte length Catches truncation that parsing will not
Canonical URL from the page The site’s own statement about which copy is authoritative

Store these beside the document, not in a separate log that will be rotated away. A corpus without per-document provenance cannot be filtered, audited or partially deleted, and all three get asked for eventually.

#Deduplication is most of the work

The web is enormously repetitive, and a training corpus that does not deduplicate is mostly boilerplate. Three layers, cheapest first:

  • Exact hash. A hash of the normalised body catches byte-identical duplicates: mirrors, syndication, the same page under two URLs. Cheap and catches a great deal.
  • URL canonicalisation. Strip tracking parameters, respect the page’s own rel="canonical", and normalise trailing slashes and case in the host. Much duplication is one document behind many addresses.
  • Near-duplicate detection. Templated pages differ only in a product name or a date. This is the expensive layer and the one that decides quality; techniques such as MinHash or SimHash exist precisely because exact hashing misses it.

Do the first two at write time. They cost almost nothing and shrink what the third layer has to examine.

#Rate, and why the address is not the fix

The instinct at scale is to add addresses. It works for per-address rate limits and for nothing else. Our measurements of what HTTP clients send by default and how their TLS handshakes differ both point the same way: most of what identifies a bulk client is produced by the client, not the network path, and a new address changes none of it.

What does help:

# One host at a time, with a delay proportional to its response time.
# A slow server is telling you something; a fixed delay ignores it.
delay = max(min_delay, last_response_seconds * courtesy_factor)

# Honour Retry-After exactly. It is standardised and it is an instruction.
if response.status == 429 and 'Retry-After' in response.headers:
    sleep(parse_retry_after(response.headers['Retry-After']))

Per-host concurrency of one or two, with backoff keyed to the host’s own latency, will collect more over a week than aggressive parallelism collects before it gets refused. Concurrency is the setting that causes most blocks, not total volume.

#Identify yourself, and mean it

A crawler’s user agent is a claim, and at this scale it is a claim you will be held to. Three positions are defensible; a fourth is not.

Position What it means in practice
A named agent with a contact URL Operators can reach you, and can allow you specifically. The only position that lets a site say yes
An honest library default Unremarkable and truthful. Fine for small, permitted work
No user agent at all Treated as more suspicious than an unrecognised one, because ordinary traffic almost always carries something
A copied browser string Contradicted by everything else your client does, and it removes any claim that access was in good faith
User-Agent: ExampleCorpBot/1.0 (+https://example.com/bot)

The fourth position is worth spelling out because it is the common one. Our measurements show a browser sending eleven headers where the median library sends four, in a different order, with none of the Sec-Fetch-* set — so a browser string over a library connection is not a disguise. It is a contradiction, and a contradiction is more distinctive than an honest default. The same holds one layer down: the TLS handshake is fixed per library and cannot be changed by anything at the HTTP layer.

If you want the option of being permitted rather than merely undetected, a named agent is the only route. It is also the only one that survives someone asking, later, what you were doing.

#Storage shape decides what you can do later

Two decisions taken on day one are expensive to reverse.

Store the raw response, not just the extracted text. Extraction rules improve, and re-extracting from stored HTML costs nothing while re-fetching a million pages costs a great deal and may no longer be possible. Compressed raw HTML is not the expensive part of a corpus.

Partition by source and by fetch date. Every operational request you will receive later — remove this domain, rebuild everything older than a date, audit what came from one site — is trivial against that layout and painful against a flat store. It costs nothing at write time.

Both follow the same principle as the provenance fields above: record what is cheap now and irrecoverable later.

#What breaks past a few million pages

These are the failures that do not appear in a pilot run.

Failure Symptom Guard
Silent content drift A parser change quietly empties a field for one site Per-source populated-field counts, compared run to run
Challenge pages stored as text Thousands of near-identical “verify you are human” documents Assert on content, not on a 200 status
Redirect loops into one page One document, thousands of URLs Hash before store; record the final URL
Encoding damage Mojibake in a fraction of documents Trust the declared charset, verify by decode
Unbounded queue growth Memory exhaustion days in Bound the frontier; persist it
Politeness state lost on restart A restart hammers hosts you had backed off from Persist per-host backoff, not just the queue

The first two are the expensive ones because they are invisible. A pipeline storing challenge pages reports a healthy success rate the whole time. This is why success measured on status codes is misleading: a 200 asserts that a response was returned, not that it was the document you asked for.

#Decide what you will not collect

Some of this is legal and jurisdiction-specific, and none of it is advice. But three categories are worth an explicit decision before the crawl starts rather than after:

  • Personal data. Obligations attach to it regardless of whether the page was public. Being able to find and delete every document from one source later is a design requirement, not a feature.
  • Material behind a login or a paywall. Public and accessible are different things, and the distinction matters more here than almost anywhere.
  • Sites that asked you not to. Whether or not the file binds you, ignoring it after reading it is a different position from never having looked.

The provenance fields above are what make any of this actionable. Without a per-document record of source and fetch time, “remove everything from that domain” is not a query you can run.

#A note on being crawled yourself

If you publish as well as collect, the same survey is worth reading from the other side. Two thirds of the sites we checked name at least one AI agent in robots.txt, and the ones writing careful rules are drawing the training-versus-answering line deliberately rather than blocking everything. Whichever side you are on, that distinction is the one worth being precise about.

#Where to go next

The mechanics of the exclusion protocol are in the robots.txt entry. Rate limiting, backoff and what Retry-After obliges you to do are covered in reading proxy and block errors. For how a bulk client is identified before it sends a header, see the two measurement studies linked above.

Frequently asked questions

Does robots.txt legally prevent me collecting training data?
No. It is a voluntary protocol standardised in RFC 9309 that records what a site operator asks crawlers to do. It is not access control and it is not a licence. What it does give you is a record of what you were told, which is a different position from never having looked, and that distinction tends to matter later.
What is the most common robots.txt parsing mistake?
Treating each User-agent line as starting a new group. Consecutive User-agent lines share the rules that follow them, and large sites rely on this heavily. In our survey of 46 sites, one groups roughly twenty agents under a single Disallow, so a parser missing the rule would score nineteen of them as unmentioned.
Should my crawler use a browser user agent?
No. A browser string over an HTTP client is contradicted by the header set, the header order and the TLS handshake, all of which we have measured. It makes the request more distinctive rather than less, and it removes any claim that access was in good faith. A named agent with a contact URL is the only position that lets an operator allow you deliberately.
Why do I need a content hash for every document?
Deduplication and change detection. The web is heavily repetitive, and a corpus that does not deduplicate is largely boilerplate. The hash also tells you on a re-crawl whether a document actually changed, which a timestamp alone cannot.
How fast can I crawl one site?
Slower than you want, and keyed to that site rather than to a global setting. Per-host concurrency of one or two with a delay proportional to the host's own response time collects more over a week than aggressive parallelism collects before being refused. Honour Retry-After exactly when you receive it.
Why store raw HTML when I only need the text?
Because extraction rules improve and re-extracting from storage costs nothing, while re-fetching a million pages costs a great deal and may no longer be permitted. Compressed raw HTML is not the expensive part of a corpus.

Sources

  1. RFC 9309: Robots Exclusion Protocol, including the grouping and matching rules
  2. RFC 9110: HTTP Semantics, status 429 and the Retry-After field
  3. RFC 6585: status code 429, Too Many Requests

Read this page as Markdown · Quote it freely under CC BY 4.0 with a link back.