Skip to content
Glossary

Web scraping

Extracting structured data from web pages automatically. Lawful in many contexts and restricted in others.

Web scraping is the automated extraction of data from websites. A scraper fetches pages and parses the response into structured records.

#The usual pipeline

  1. Fetch — retrieve the page, often through a proxy.
  2. Render — only if the content requires JavaScript; a headless browser is expensive and often unnecessary.
  3. Parse — extract fields from the markup or an underlying JSON response.
  4. Validate — confirm you received real content rather than a block page.
  5. Store — persist, deduplicate, and record when each record was captured.

#The step most often skipped

Validation. Many anti-bot systems answer with 200 and an empty or decoy payload. A pipeline that checks only the status code will happily record thousands of empty results and report a high success rate while collecting nothing.

Legality depends on jurisdiction, on what is collected, on whether access controls are bypassed, and on what is done with the result afterwards. Personal data brings data-protection law into scope regardless of how public the page is. This is not legal advice.

#Where a pipeline usually breaks

Collection is the visible step and rarely the one that fails silently. The two that do are validation and change detection.

Step Typical silent failure
Fetch A 200 carrying a challenge page instead of content
Parse A changed selector yields empty fields, not an error
Validate Skipped entirely, so empty records are stored
Store Yesterday’s good data overwritten by today’s empty run

Every row produces a job that reports success. This is why success measured on status codes is misleading: the pipeline is healthy by that definition while collecting nothing.

#The check that catches all four

Assert on content, not on transport. Require that a field you actually need is present and plausible before a record is accepted, and compare the day’s record count against the previous run. A sudden collapse in populated fields is the signal that something upstream changed, and it arrives days before anyone notices the data is wrong.

#Before you start

  • Look for an API first. It is cheaper, more stable and usually permitted.
  • Read robots.txt and the terms of use. They state the operator’s position.
  • Consider what you are collecting. Personal data carries obligations independent of how you obtained it.
  • Pace the work. Concurrency causes more blocks than volume does.

Legality varies by jurisdiction, by what is collected and by how access is obtained. Take advice for anything at scale rather than relying on a general summary.

Frequently asked questions

How do I stop a scraper silently collecting nothing?
Validate content rather than status codes. Require a field you actually need to be present and plausible before accepting a record, and compare each run's populated-field counts against the previous run.
Is web scraping legal?
It depends on the jurisdiction, on what you collect and on how you obtain access. Public data, personal data and material behind a login are treated very differently. Take advice for anything at scale rather than relying on a general answer.
Should I use an API instead?
Where one exists, almost always. It is cheaper to run, far more stable than parsing HTML, and usually explicitly permitted. Check the network panel of the page you are targeting, because many sites fetch their content from an API you can call directly.

Sources

  1. RFC 9309: Robots Exclusion Protocol
  2. RFC 9110: HTTP Semantics

Related terms