---
title: "Reverse proxy"
url: https://proxy.wiki/glossary/reverse-proxy/
type: Glossary Term
author: "proxy.wiki editorial"
published: 2026-08-19
updated: 2026-08-29
site: proxy.wiki
topics: ["Proxy fundamentals"]
license: CC BY 4.0 — quote freely with attribution to https://proxy.wiki/
---

# Reverse proxy

> A proxy deployed in front of a server. Clients connect to it as if it were the origin, and it forwards to the real backend.

**A reverse proxy sits in front of one or more origin servers and accepts connections on their behalf.** The client believes it is talking to the real service; the reverse proxy decides which backend actually answers.

## What it is used for

- **Load balancing** across several backends.

- **TLS termination**, so certificates are managed in one place.

- **Caching** of static responses.

- **Filtering**, including the [anti-bot systems](/glossary/anti-bot-system/) that scrapers run into.

Nginx, HAProxy, Cloudflare and most CDNs operate as reverse proxies.

## Why it matters when scraping

The thing blocking your request is very often a reverse proxy rather than the application itself. That is why a block can arrive in single-digit milliseconds and why the response frequently looks nothing like the site’s normal error page.

## Commonly confused with

A [forward proxy](/glossary/forward-proxy/) represents the client. A reverse proxy represents the server. The word “proxy” alone almost always means the forward kind.

## What the scraper actually meets

Almost every large site answers from a reverse proxy rather than from the application. That intermediary, not the application, usually issues the block. It explains several behaviours that otherwise look inconsistent.

- **The response arrives too quickly to have reached the application.** A block served in a few milliseconds came from the edge.

- **The error page does not match the site’s design.** Edge providers serve their own challenge pages.

- **Headers name the intermediary.** `Server`, `Via` and vendor-specific ray or trace identifiers appear on the response.

This matters when you diagnose a failure. A [CAPTCHA](/glossary/captcha/) or a 403 from the edge means the request never reached the origin, so changing what you ask for will not help. Changing how the request looks might.

## Inspecting the layer

```
curl -sI https://example.com/ | grep -iE 'server|via|cf-|x-cache|x-served-by'
```

Named intermediaries and cache-status headers tell you an edge is present. Absence proves nothing, because many operators strip these headers deliberately.

## Frequently asked questions

### Why does a reverse proxy matter if I am only reading a public page?

Because it decides whether your request reaches the site at all. Rate limits, geographic rules and bot challenges are usually enforced at that layer, so the block you receive often has nothing to do with the page you asked for.

### Can I tell which reverse proxy a site uses?

Sometimes. Response headers such as Server and Via, and vendor trace identifiers, name it. Many operators remove these, so a missing header is not evidence of a missing proxy.

### Is a load balancer a reverse proxy?

In practice they overlap. A load balancer distributes connections across backends; a reverse proxy terminates the connection and may also cache, filter and rewrite. Most modern edge products do both.

## Sources

1. [RFC 9110: HTTP Semantics, the Via header field](https://www.rfc-editor.org/rfc/rfc9110.html)
