Web scraping
Extracting structured data from web pages automatically. Lawful in many contexts and restricted in others.
Web scraping is the automated extraction of data from websites. A scraper fetches pages and parses the response into structured records.
#The usual pipeline
- Fetch — retrieve the page, often through a proxy.
- Render — only if the content requires JavaScript; a headless browser is expensive and often unnecessary.
- Parse — extract fields from the markup or an underlying JSON response.
- Validate — confirm you received real content rather than a block page.
- Store — persist, deduplicate, and record when each record was captured.
#The step most often skipped
Validation. Many anti-bot systems answer with 200 and an empty or decoy payload. A pipeline that checks only the status code will happily record thousands of empty results and report a high success rate while collecting nothing.
#Legal note
Legality depends on jurisdiction, on what is collected, on whether access controls are bypassed, and on what is done with the result afterwards. Personal data brings data-protection law into scope regardless of how public the page is. This is not legal advice.
#Where a pipeline usually breaks
Collection is the visible step and rarely the one that fails silently. The two that do are validation and change detection.
| Step | Typical silent failure |
|---|---|
| Fetch | A 200 carrying a challenge page instead of content |
| Parse | A changed selector yields empty fields, not an error |
| Validate | Skipped entirely, so empty records are stored |
| Store | Yesterday’s good data overwritten by today’s empty run |
Every row produces a job that reports success. This is why success measured on status codes is misleading: the pipeline is healthy by that definition while collecting nothing.
#The check that catches all four
Assert on content, not on transport. Require that a field you actually need is present and plausible before a record is accepted, and compare the day’s record count against the previous run. A sudden collapse in populated fields is the signal that something upstream changed, and it arrives days before anyone notices the data is wrong.
#Before you start
- Look for an API first. It is cheaper, more stable and usually permitted.
- Read robots.txt and the terms of use. They state the operator’s position.
- Consider what you are collecting. Personal data carries obligations independent of how you obtained it.
- Pace the work. Concurrency causes more blocks than volume does.
Legality varies by jurisdiction, by what is collected and by how access is obtained. Take advice for anything at scale rather than relying on a general summary.