Skip to content
Topic

Web scraping

Collecting data from the web at scale: client configuration, reading the failures correctly, and the silent bug that returns nothing.

Most scraping problems are not scraping problems. They are HTTP problems, session problems or measurement problems wearing a different hat, and they are much easier to fix once named correctly. That is what this section is organised around.

Start with the client

Whatever language you use, the first thing to get right is the request itself. Each of these guides was written by running the commands and recording what happened, not by summarising documentation:

  • Proxies with curl — the fastest way to prove an endpoint works at all, and the socks5h trap that fails silently.
  • Proxies in Python — requests, httpx and aiohttp, with session reuse and the TLS caveat.
  • Proxies in Node.js — the runtime that ignores the proxy environment variables entirely.
  • The environment variables — why HTTP_PROXY in uppercase is the one spelling curl deliberately refuses.

Then read the failures correctly

A block tells you which layer refused you, and that decides what to change. Changing your address when the refusal came from the transport layer wastes money and time.

  • 403, 407 and 429 — three refusals with three different meanings and three different fixes.
  • Rate limiting — what Retry-After obliges you to do, and why rotating around a limit often fails.
  • Anti-bot systems — the layers a request passes through, and which of them a new address actually affects.

The failure nobody notices

The most expensive scraping bug is not a block. It is a job that reports success while collecting nothing.

A challenge page returns 200. A consent wall returns 200. A page served to a suspected bot with the data stripped out returns 200. Any pipeline that measures health by status code is blind to all three, and it stays blind until somebody opens the output weeks later.

The fix is to assert on content. Require a field you actually need to be present and plausible before a record is accepted, and compare each run's populated-field counts against the previous one. This is covered in success rate, which also explains why published success figures are usually not comparable with each other.

Pacing and etiquette

Concurrency causes more blocks than total volume does, because parallelism turns an unremarkable request rate into a burst. Raise it in steps and watch the measured success rate rather than the throughput; the right setting is the highest one at which quality has not started to fall.

robots.txt is worth reading before a large crawl. It is advisory rather than enforceable, standardised in RFC 9309, and it records what the site operator asks of crawlers. Ignoring it rarely blocks you and does remove any claim that you were not told.

Browsers, and when you do not need one

A headless browser costs orders of magnitude more than an HTTP request, so establish that you need one first. Open the network panel and watch what the page fetches: sites that appear to require rendering very often pull their content from a single API endpoint you can call directly.

Where you do need one, note that its defaults announce it. Automation flags, default window sizes and missing plugin lists are all readable by the page, and a browser left at its defaults is easier to identify than a plain HTTP client.

In this section

12 pages
Glossary Term

Concurrency

How many proxy connections you may run at once — often the real constraint on throughput, not bandwidth.

2 min read

Glossary Term

Headless browser

A real browser engine driven programmatically with no visible window. Powerful, slow and expensive in bandwidth.

2 min read

Guide

How to use a proxy with curl

Every curl proxy flag that matters, the socks5h trap, how to verify your exit address, and how to read the errors.

8 min read

Guide

Proxies for git, npm and pip

git config https.proxy is silently ignored, npm's equivalent works, and pip blames the package. Every setting tested against a closed port.

8 min read

Glossary Term

Rate limiting

Capping how many requests a client may make in a period. The 429 status code is defined in RFC 6585.

2 min read

Glossary Term

robots.txt

A file telling crawlers which paths a site prefers they avoid. A convention, not an access control.

2 min read

Glossary Term

Success rate

The share of requests that return a usable response. Meaningless without a sample size and a named target.

2 min read

Glossary Term

User agent

The header a client uses to identify itself. Trivially changed, and therefore weak evidence on its own.

2 min read

Glossary Term

Web scraping

Extracting structured data from web pages automatically. Lawful in many contexts and restricted in others.

2 min read