---
title: "Web scraping"
url: https://proxy.wiki/glossary/web-scraping/
type: Glossary Term
author: "proxy.wiki editorial"
published: 2026-08-19
updated: 2026-08-29
site: proxy.wiki
topics: ["Web scraping"]
license: CC BY 4.0 — quote freely with attribution to https://proxy.wiki/
---

# Web scraping

> Extracting structured data from web pages automatically. Lawful in many contexts and restricted in others.

**Web scraping is the automated extraction of data from websites.** A scraper fetches pages and parses the response into structured records.

## The usual pipeline

- **Fetch** — retrieve the page, often through a [proxy](/glossary/proxy-server/).

- **Render** — only if the content requires JavaScript; a [headless browser](/glossary/headless-browser/) is expensive and often unnecessary.

- **Parse** — extract fields from the markup or an underlying JSON response.

- **Validate** — confirm you received real content rather than a block page.

- **Store** — persist, deduplicate, and record when each record was captured.

## The step most often skipped

Validation. Many [anti-bot systems](/glossary/anti-bot-system/) answer with 200 and an empty or decoy payload. A pipeline that checks only the status code will happily record thousands of empty results and report a high [success rate](/glossary/success-rate/) while collecting nothing.

## Legal note

Legality depends on jurisdiction, on what is collected, on whether access controls are bypassed, and on what is done with the result afterwards. Personal data brings data-protection law into scope regardless of how public the page is. This is not legal advice.

## Where a pipeline usually breaks

Collection is the visible step and rarely the one that fails silently. The two that do are validation and change detection.

| Step | Typical silent failure |
| --- | --- |
| Fetch | A 200 carrying a challenge page instead of content |
| Parse | A changed selector yields empty fields, not an error |
| Validate | Skipped entirely, so empty records are stored |
| Store | Yesterday’s good data overwritten by today’s empty run |

Every row produces a job that reports success. This is why [success measured on status codes](/glossary/success-rate/) is misleading: the pipeline is healthy by that definition while collecting nothing.

## The check that catches all four

Assert on content, not on transport. Require that a field you actually need is present and plausible before a record is accepted, and compare the day’s record count against the previous run. A sudden collapse in populated fields is the signal that something upstream changed, and it arrives days before anyone notices the data is wrong.

## Before you start

- **Look for an API first.** It is cheaper, more stable and usually permitted.

- **Read [robots.txt](/glossary/robots-txt/) and the terms of use.** They state the operator’s position.

- **Consider what you are collecting.** Personal data carries obligations independent of how you obtained it.

- **Pace the work.** [Concurrency](/glossary/concurrency/) causes more blocks than volume does.

Legality varies by jurisdiction, by what is collected and by how access is obtained. Take advice for anything at scale rather than relying on a general summary.

## Frequently asked questions

### How do I stop a scraper silently collecting nothing?

Validate content rather than status codes. Require a field you actually need to be present and plausible before accepting a record, and compare each run's populated-field counts against the previous run.

### Is web scraping legal?

It depends on the jurisdiction, on what you collect and on how you obtain access. Public data, personal data and material behind a login are treated very differently. Take advice for anything at scale rather than relying on a general answer.

### Should I use an API instead?

Where one exists, almost always. It is cheaper to run, far more stable than parsing HTML, and usually explicitly permitted. Check the network panel of the page you are targeting, because many sites fetch their content from an API you can call directly.

## Sources

1. [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)
2. [RFC 9110: HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110.html)
