Concurrency
How many proxy connections you may run at once — often the real constraint on throughput, not bandwidth.
Collecting data from the web at scale: client configuration, reading the failures correctly, and the silent bug that returns nothing.
Most scraping problems are not scraping problems. They are HTTP problems, session problems or measurement problems wearing a different hat, and they are much easier to fix once named correctly. That is what this section is organised around.
Whatever language you use, the first thing to get right is the request itself. Each of these guides was written by running the commands and recording what happened, not by summarising documentation:
socks5h trap that fails silently.HTTP_PROXY in uppercase is the one spelling curl deliberately refuses.A block tells you which layer refused you, and that decides what to change. Changing your address when the refusal came from the transport layer wastes money and time.
Retry-After obliges you to do, and why rotating around a limit often fails.The most expensive scraping bug is not a block. It is a job that reports success while collecting nothing.
A challenge page returns 200. A consent wall returns 200. A page served to a suspected bot with the data stripped out returns 200. Any pipeline that measures health by status code is blind to all three, and it stays blind until somebody opens the output weeks later.
The fix is to assert on content. Require a field you actually need to be present and plausible before a record is accepted, and compare each run's populated-field counts against the previous one. This is covered in success rate, which also explains why published success figures are usually not comparable with each other.
Concurrency causes more blocks than total volume does, because parallelism turns an unremarkable request rate into a burst. Raise it in steps and watch the measured success rate rather than the throughput; the right setting is the highest one at which quality has not started to fall.
robots.txt is worth reading before a large crawl. It is advisory rather than enforceable, standardised in RFC 9309, and it records what the site operator asks of crawlers. Ignoring it rarely blocks you and does remove any claim that you were not told.
A headless browser costs orders of magnitude more than an HTTP request, so establish that you need one first. Open the network panel and watch what the page fetches: sites that appear to require rendering very often pull their content from a single API endpoint you can call directly.
Where you do need one, note that its defaults announce it. Automation flags, default window sizes and missing plugin lists are all readable by the page, and a browser left at its defaults is easier to identify than a plain HTTP client.
How many proxy connections you may run at once — often the real constraint on throughput, not bandwidth.
A real browser engine driven programmatically with no visible window. Powerful, slow and expensive in bandwidth.
Every curl proxy flag that matters, the socks5h trap, how to verify your exit address, and how to read the errors.
git config https.proxy is silently ignored, npm's equivalent works, and pip blames the package. Every setting tested against a closed port.
Capping how many requests a client may make in a period. The 429 status code is defined in RFC 6585.
A file telling crawlers which paths a site prefers they avoid. A convention, not an access control.
The share of requests that return a usable response. Meaningless without a sample size and a named target.
The header a client uses to identify itself. Trivially changed, and therefore weak evidence on its own.
Node ignores the proxy environment variables entirely. Working configuration for every client, tested against a proxy that logs each request.
Working proxy configuration for the three main Python HTTP clients, with session reuse, retries and the TLS caveat.
Extracting structured data from web pages automatically. Lawful in many contexts and restricted in others.
46 sites, 18 AI agents, read from robots.txt. ClaudeBot was blocked by 22 and GPTBot by 16, and the ordering is strict.