Crawl budget
Crawl budget is the informal name for how much of a site a crawler will fetch in a given period.
Crawl budget is the informal name for how much of a site a crawler will fetch in a given period. No specification defines it. The term comes from search-engine documentation, where it describes the interaction between what a server tolerates and what the crawler judges worth fetching.
#What actually constrains a crawl
- The server’s response. Slow responses and errors cause a well-behaved crawler to reduce its rate.
- The crawler’s own judgement of how much of the site is worth revisiting, which no site operator can observe directly.
- The rules in robots.txt, which remove paths from consideration entirely.
Nobody outside the search engine can state the size of the budget for a given site, and we do not publish a figure for it. What can be measured is your own server logs: which URLs were fetched, how often, and what they returned.
#What the robots standard does and does not cover
RFC 9309 standardises the user-agent grouping and the allow and disallow rules, and sets two useful constants: a parser must accept at least 500 kibibytes of the file, and a cached copy should not be used for more than 24 hours unless caching directives say otherwise. It does not define Crawl-delay. That directive is a widely implemented extension, honoured by some crawlers and ignored by others, so do not rely on it to control anyone.
#Spending it badly
Most waste comes from generating unbounded URL spaces: faceted filters in every combination, session identifiers in paths, infinite calendars, and sorting parameters that produce a new URL for the same content. Each is a set of distinct URLs serving duplicate content, and a crawler will happily consume all of it.
#Commonly confused with
Crawl rate is a speed, measured at a moment. Crawl budget is a total, spent over a period. A site can be crawled slowly and thoroughly or quickly and shallowly, so knowing one tells you little about the other. Neither is the Crawl-delay directive, which is a request a crawler may decline.
The same reasoning applies from the other side. Fetching the same content under many URLs wastes your bandwidth and your concurrency on nothing. Normalise URLs before queueing them, deduplicate on the normalised form, and cap the depth of parameter combinations you are prepared to follow. See also honeypot links, which exist to consume exactly this budget.