robots.txt
A file telling crawlers which paths a site prefers they avoid. A convention, not an access control.
robots.txt is a plain-text file at the root of a domain that states which paths automated clients are asked not to fetch. It is advisory: it expresses a preference and enforces nothing.
#How it reads
User-agent: *
Disallow: /admin/
Allow: /
Sitemap: https://example.com/sitemap.xml
Directives are grouped by user agent, with * as the fallback. The Sitemap line is how crawlers discover a site’s URL index.
#What it is not
- Not security. Listing a path here advertises it. Anything sensitive needs authentication.
- Not enforcement. Compliance is voluntary. Major search engines honour it; many other clients do not.
- Not a licence. Being permitted to fetch a page says nothing about your right to reuse its contents.
#Why it matters when scraping
Ignoring it is technically trivial and reputationally expensive. It is frequently cited in disputes as evidence of intent, and it is often the first thing examined when a site owner objects.
#What the standard actually says
RFC 9309 published the Robots Exclusion Protocol as a standard in 2022, after decades as a convention. It defines the file location, the grammar and the matching rules. It does not make the file enforceable; compliance remains voluntary.
| Directive | Meaning |
|---|---|
User-agent |
Which crawler the following rules apply to |
Disallow |
A path prefix the crawler should not fetch |
Allow |
An exception within a disallowed prefix |
Sitemap |
Where the sitemap lives. Independent of the groups |
The most specific matching group wins, and the longest matching rule within it wins. A crawler that matches no named group uses the * group if one exists.
#What it is not
- Not access control. A disallowed path is still served to anyone who requests it.
- Not a privacy measure. The file is public and lists paths the operator considers sensitive.
- Not a legal boundary in itself. It is evidence of the operator’s stated wishes, which is a different thing.
#Why it matters when scraping
Ignoring the file rarely blocks you, because it is advisory. It does establish that you were told, and it removes any claim that access was unintentional. Read it before a large crawl, honour the crawl-rate signals where present, and treat it as the operator’s stated position rather than as an obstacle. See web scraping.