A retrieval corpus that is a month old does not fail loudly. It answers fluently, cites a URL, and gives you last month’s facts with no hedging anywhere in the output. That asymmetry is the whole reason fetching for retrieval-augmented generation is a different job from crawling. A crawler that misses a page has a smaller corpus. A RAG index that holds a stale page has a confident wrong answer waiting for the first question that touches it.
Most RAG fetchers are crawlers with the crawl loop deleted. That inheritance is wrong in almost every detail: the URL set, the refresh policy, where politeness has to be spent, and above all what counts as a successful fetch. This guide covers those differences and the failures they produce. It is the companion to collecting training data at scale, which is the opposite problem and takes the opposite advice.
#Staleness here is a wrong answer, not a gap
A training corpus tolerates age. The model is frozen at a cut-off regardless, nobody expects the weights to know yesterday, and a document from last year is still a valid example of language and of most facts. Retrieval is sold on exactly the property that training data does not have: that it is current. When it is not, nothing in the system notices.
The two failure modes are not equally bad, and this is the point people get backwards. If retrieval returns nothing, a well-built system says it does not know, and the user goes and looks. If retrieval returns a document that was correct in March, the system answers, attaches the link, and the link makes the answer look checked. A user who follows it finds a page that now says something else. The citation is then arguing against the answer it was attached to.
The practical consequence is small and almost always skipped: the fetch time has to travel with the text into the prompt, not sit in a metadata table for your dashboard. A model that can see retrieved 2026-08-02 above a chunk can hedge. One that cannot, will not.
# The retrieval unit is text plus provenance, not text.
chunk = (
f"Source: {doc.final_url}\n"
f"Fetched: {doc.fetched_at:%Y-%m-%d %H:%M} UTC\n"
f"Section: {' > '.join(chunk.heading_path)}\n\n"
f"{chunk.text}"
)
Embedding the header along with the body slightly pollutes the vector, so most people strip it and lose the provenance. Keep it out of the embedded text and prepend it at prompt-assembly time instead. The retrieval scores stay clean and the model still sees the date.
#Live fetch or a pre-built index: choose a failure mode
The trade-off is usually presented as cost. Cost is the axis that changes least between the two designs. What actually differs is latency, and what the system does when the fetch does not work.
| Live fetch at query time | Pre-built index | |
|---|---|---|
| User waits for | The slowest host in the set, plus extraction | A vector lookup |
| On failure | Nothing retrieved | Last stored copy, of unknown age |
| Is the failure visible? | Yes — a timeout or a status code | No — the answer looks normal |
| Coverage | Only what is reachable right now | Everything ever fetched, including pages since removed |
| Blocked or rate-limited source | Surfaces immediately as a user-visible error | Surfaces months later as a quietly frozen document |
| Extraction bugs | Hit one query | Hit every query until someone notices |
Neither column is the safe one. A live fetch fails loudly and leaves the model with nothing to ground on, which at least degrades honestly. An index fails silently and keeps answering. Most production systems end up hybrid: retrieve from the index, and fetch live only the small number of URLs the query strongly implicates.
If you build the hybrid, make one rule non-negotiable. When the live fetch fails and you substitute the stored copy, that substitution has to be labelled in the context you hand the model. Silently falling back to the index converts a visible failure into an invisible one, which is the exact trade you were trying to avoid.
#You are refetching a known set, not discovering an unknown one
A crawler’s URL set grows. It starts from seeds, extracts links, and the frontier is the data structure the whole design is built around. A RAG source list is enumerated by a person or read from a sitemap, and it barely moves: a documentation tree, a changelog, a pricing page, a policy page, a status feed, a handful of reference articles. Where a crawler is a discovery problem, this is a scheduling problem.
That inversion changes what to optimise. Address diversity, the default answer to every scraping question, matters much less here — you are not trying to look like a thousand visitors. Per-host politeness matters much more, because a crawler spreads its request budget across thousands of hosts while a refresher aims the same budget at ten. The same total volume is a completely different per-host rate, and per-host rate is what gets refused. Rate limiting and concurrency are the settings that decide whether this works.
Being one identifiable, predictable client is a better position than being many anonymous ones. A stable user-agent with a contact URL, a fixed schedule and honoured Retry-After gives a source operator a reason not to block you, and gives you someone to email when they do. Check what your HTTP client sends before you configure anything — our measurements of default request headers show how little most clients volunteer. Read the source’s robots.txt once per source and store the verdict, rather than fetching it on every pass.
Proxies still earn their place here, for a different reason than in bulk crawling: reachability and geography. Where a page differs by country — pricing, availability, legal text — you need an exit in each country you claim to cover, and you need to record which exit produced each chunk. A corpus that mixes country variants of the same URL without labelling them will answer pricing questions with whichever variant embedded closest.
Schedule per source, not globally. The interval is a property of the source, not of your pipeline.
| Source class | What sets the interval | Best signal to key off |
|---|---|---|
| Reference docs | The product’s release cadence | Changelog or release feed, not the doc page |
| Pricing and terms | Rare, irregular, high impact when missed | Conditional request plus a change alert to a human |
| Status and incidents | Continuous during an incident, idle otherwise | The provider’s feed or webhook if one exists |
| Blog or news index | Publication cadence | RSS or Atom where published |
| Community wiki | Continuous, uneven | Last-modified from the site’s API where offered |
#Revalidate rather than refetch
Conditional requests are defined in RFC 9110 and the caching rules around them in RFC 9111. They are more useful in a RAG refresher than in a crawler, because a crawler mostly sees each URL once while a refresher sees the same URL over and over.
Run against the RFC 9111 page itself, on 6 September 2026, with the ETag the server had just issued:
U=https://www.rfc-editor.org/rfc/rfc9111.html
ET=$(curl -sI "$U" | tr -d '\r' | awk -F': ' '/^etag:/{print $2}')
echo "$ET"
# W/"67c426a73321f73db9ffdf2508376ec5"
curl -s -o /dev/null -w '%{http_code} %{size_download}\n' "$U"
# 200 232630
curl -s -o /dev/null -H "If-None-Match: $ET" -w '%{http_code} %{size_download}\n' "$U"
# 304 0
The bytes are the least interesting saving. What a 304 buys you is skipping the parse, the extraction, the chunking, the embedding call and the index write — the expensive half of the pipeline, all of it downstream of a document that did not change. Note the W/ prefix: that is a weak validator, meaning semantically equivalent rather than byte-identical, which is exactly the comparison you want here.
Keep two independent change signals per URL: the server’s opinion, and yours. The server’s is the ETag or Last-Modified. Yours is a hash of the extracted text, after boilerplate removal. They disagree usefully. Pages with a timestamped footer or a rotating advert issue a new ETag on every request while hashing identically after extraction; you want the second signal to stop you re-embedding those. The reverse case is rarer and worth an alert.
import hashlib
def refresh(url, state, client):
headers = {}
if state.get("etag"):
headers["If-None-Match"] = state["etag"]
if state.get("last_modified"):
headers["If-Modified-Since"] = state["last_modified"]
r = client.get(url, headers=headers, timeout=20.0)
state["checked_at"] = utcnow()
if r.status_code == 304:
return "unchanged"
if r.status_code != 200:
# Never overwrite the stored copy with an error page.
state["degraded"] = r.status_code
return f"error:{r.status_code}"
state["etag"] = r.headers.get("ETag")
state["last_modified"] = r.headers.get("Last-Modified")
state["final_url"] = str(r.url)
text = extract(r.text) # boilerplate already stripped
digest = hashlib.sha256(text.encode()).hexdigest()
if digest == state.get("digest"):
return "unchanged-after-extraction" # validator moved, content did not
state["digest"] = digest
return store(url, text, state)
The line that matters most is the one that refuses to overwrite on a non-200. A refresh loop that writes whatever it received is how a working corpus turns into a corpus of error pages over a weekend, and every document keeps a fresh-looking timestamp while it happens. Mark the source degraded, keep the last good copy, and let the answering layer decide whether to use an old document or say nothing.
#Extraction quality is answer quality
Whatever your extractor emits becomes a chunk. Whatever gets retrieved becomes a citation. There is no stage in between where a human reads it. Boilerplate is therefore not cosmetic damage — it is content your system will quote.
Two distinct harms. The first is dilution: navigation appears on every page of a host, so after chunking, the most duplicated text in a single-site corpus is the menu. Short queries match short repetitive chunks. The second is attribution: the user clicks the citation and cannot find the sentence on the page, because the sentence came from a consent dialog that no longer renders for them.
| What leaks in | How it shows up in retrieval | Cheap detection |
|---|---|---|
| Cookie and consent text | Chunks about privacy choices retrieved for unrelated queries | Identical text across unrelated hosts |
| Navigation and footer | The most-duplicated text in the corpus | Per-host line frequency; near-universal lines are chrome |
| “Related articles” blocks | Citations to pages you never fetched | Link density per block |
| Relative dates (“updated 3 days ago”) | Permanently wrong once stored | Resolve against the fetch timestamp at extraction time |
| Flattened tables | Rows merged into one unreadable line; values pair with the wrong label | Compare cell count before and after extraction |
| Empty single-page-app shell | A short chunk of nothing, retrieved for its host | Extracted length far below the page’s own byte length |
Three habits handle most of it. Deduplicate across pages of the same host before chunking — you do not need a boilerplate model, a per-host line frequency count finds the chrome. Keep the heading path with each chunk; it is the cheapest relevance signal available and it survives chunking intact. And store the raw response alongside the extracted text, because re-extraction is the only cheap way to fix an extractor bug retroactively. Without the raw copy, every improvement to your extractor costs a full refetch of every source.
#A 200 that is not the document
Challenge pages, consent walls, login interstitials, geo notices and empty app shells all return 200. Each one is a perfectly valid HTTP response containing the wrong document, and every part of a naive pipeline treats it as a success: the status check passes, the extractor produces text, the embedder embeds it, the index stores it. This is why success rate measured on status codes is the wrong metric — a 200 asserts that a response arrived, not that it is what you asked for. If the response is a genuine refusal, 403, 407 and 429 mean three different things and only one of them is about your proxy.
What makes this worse in RAG than in bulk collection is arithmetic. A crawl of millions dilutes a poisoned source into noise. A corpus of a few hundred documents may have one source for a topic, and if that source is poisoned then every query about the topic retrieves a consent banner. The retrieval layer will not object. “Verify you are human” embeds like any other sentence.
The fix is to assert positively, per source. You know what that page should contain, so check for it rather than checking a status code.
# Per-source contract, checked on every fetch.
SOURCES = {
"vendor-pricing": {
"url": "https://example.com/pricing",
"must_contain": ["Plans", "per month"], # stable page furniture
"min_chars": 1200, # set from your own history
},
}
def verify(source, text):
spec = SOURCES[source]
missing = [m for m in spec["must_contain"] if m not in text]
if missing:
raise SourceShapeChanged(f"{source}: missing {missing}")
if len(text) < spec["min_chars"]:
raise SourceShapeChanged(f"{source}: {len(text)} chars, expected more")
The min_chars value is a configuration choice, not a measurement — derive it from what that source has produced historically, and treat any number written in from memory as a guess. A phrase blocklist of known interstitial text is worth having as a backstop, but only as a backstop: the phrases change, and a blocklist is the check that silently stops working.
One caution on the obvious escalation. Switching a blocked source to a headless browser changes what you store as well as whether you get a response, because you are now capturing post-JavaScript DOM rather than served HTML. Re-verify extraction on that source after the switch instead of assuming it improved.
#Detecting that a source has changed shape
Silent parser failure is the most expensive bug in this class of system, and it is more expensive here than in bulk crawling for the same reason as above: the corpus is small, so one broken source is a large fraction of it. A single site quietly returning empty extractions in a crawl of millions is a rounding error. In a ten-source RAG index it may be the only thing that could have answered the question.
Aggregate monitoring cannot see it. Nine healthy sources and one returning nothing barely moves an aggregate. The comparison has to be per source, against that source’s own previous run.
# Compare this run to the last one. Per source, never aggregated.
SHRINK = 0.5 # a starting point; set it from your own run history
for src, now in this_run.items():
prev = last_run.get(src)
if now["chars"] == 0:
alert(f"{src}: extraction empty, HTTP {now['status']}")
elif now["final_host"] != now["requested_host"]:
alert(f"{src}: redirected off-host to {now['final_host']}")
elif prev and now["chars"] < prev["chars"] * SHRINK:
alert(f"{src}: text shrank {prev['chars']} -> {now['chars']}")
elif prev and now["digest"] == prev["digest"] and src in EXPECTED_DAILY:
alert(f"{src}: unchanged since {prev['checked_at']}, expected daily change")
The last branch catches the failure nobody instruments: a source that stopped changing. A frozen upstream, a cache stuck in front of it, or a redirect to an archived copy all produce a document that looks perfectly healthy and never updates again. In a system whose entire value is freshness, that is worth an alert as much as an empty extraction is.
Alerting is not enough on its own. A source that fails its contract should be quarantined, and quarantine here means four separate things:
- Retain the old copy. Deleting it turns a known-stale answer into a silent gap, which is harder to notice.
- Mark it stale, carrying the timestamp of the last fetch that actually succeeded.
- Tell the answering layer, so the model can qualify the answer or decline rather than cite it flatly.
- Exclude the chunk from retrieval, or label it explicitly where it lands in the context window.
Continuing to serve chunks from a source you know is broken, because deleting them would leave a hole, is how a small bug becomes a wrong answer with a citation attached.
#What this does not cover
The only figures on this page are protocol status codes and the two byte counts printed by the curl commands shown: a full GET of the RFC 9111 page returned 200 with 232,630 bytes downloaded, and the same request carrying the server’s ETag returned 304 with an empty body. Both ran on 6 September 2026 against one server. That demonstrates the mechanism. It says nothing about how many servers support conditional requests, and we have not measured that.
We have not benchmarked any extraction library against any site, so there is no recommendation here between readability-style extractors, a purpose-built parser or a rendered DOM. We have not measured how often a consent wall or challenge page returns 200 in practice, only that it does. We have not tested the conditional-request path through a forward proxy, and some intermediaries rewrite or strip validators — verify that against your own gateway rather than assuming the 304 survives.
Chunking strategy, embedding model choice, retrieval scoring and reranking are all out of scope. This guide is about getting bytes that deserve to be retrieved; what happens to them afterwards is a separate set of decisions with separate failure modes.
Nothing here is legal advice, and nothing here is a claim about any proxy provider’s performance. Whether you may store and re-serve a source’s text is a licensing question this page does not answer: robots.txt records what a crawler was asked to do, and it does not settle reuse. Where the answer matters commercially, it is worth getting in writing from the source rather than inferring it from a file.