Skip to content
Glossary

Crawl budget

Crawl budget is the informal name for how much of a site a crawler will fetch in a given period.

Crawl budget is the informal name for how much of a site a crawler will fetch in a given period. No specification defines it. The term comes from search-engine documentation, where it describes the interaction between what a server tolerates and what the crawler judges worth fetching.

#What actually constrains a crawl

  • The server’s response. Slow responses and errors cause a well-behaved crawler to reduce its rate.
  • The crawler’s own judgement of how much of the site is worth revisiting, which no site operator can observe directly.
  • The rules in robots.txt, which remove paths from consideration entirely.

Nobody outside the search engine can state the size of the budget for a given site, and we do not publish a figure for it. What can be measured is your own server logs: which URLs were fetched, how often, and what they returned.

#What the robots standard does and does not cover

RFC 9309 standardises the user-agent grouping and the allow and disallow rules, and sets two useful constants: a parser must accept at least 500 kibibytes of the file, and a cached copy should not be used for more than 24 hours unless caching directives say otherwise. It does not define Crawl-delay. That directive is a widely implemented extension, honoured by some crawlers and ignored by others, so do not rely on it to control anyone.

#Spending it badly

Most waste comes from generating unbounded URL spaces: faceted filters in every combination, session identifiers in paths, infinite calendars, and sorting parameters that produce a new URL for the same content. Each is a set of distinct URLs serving duplicate content, and a crawler will happily consume all of it.

#Commonly confused with

Crawl rate is a speed, measured at a moment. Crawl budget is a total, spent over a period. A site can be crawled slowly and thoroughly or quickly and shallowly, so knowing one tells you little about the other. Neither is the Crawl-delay directive, which is a request a crawler may decline.

The same reasoning applies from the other side. Fetching the same content under many URLs wastes your bandwidth and your concurrency on nothing. Normalise URLs before queueing them, deduplicate on the normalised form, and cap the depth of parameter combinations you are prepared to follow. See also honeypot links, which exist to consume exactly this budget.

Frequently asked questions

Is crawl budget a real, published number?
No. It is a description of behaviour, not a quantity any search engine publishes per site. Anyone quoting a specific figure for your site is guessing. What you can measure is your own server log: which URLs were requested by which crawler, how often, and what each returned.
Does Crawl-delay in robots.txt control search engine crawlers?
Not reliably. RFC 9309 standardises user-agent grouping with allow and disallow rules and does not define Crawl-delay at all. It is an extension that some crawlers honour and others ignore, so treat it as a request rather than a control and use server-side limits when you need certainty.
What wastes crawl budget most?
Unbounded URL spaces that serve duplicate content: every combination of faceted filters, sort parameters, session identifiers in paths, and calendars with no end date. Each generates distinct URLs for the same page. Canonical URLs, disallow rules and parameter handling all reduce the waste.

Sources

  1. RFC 9309: the Robots Exclusion Protocol, including parse limits and caching
  2. Google Search Central: crawl budget management for large sites

Related terms