A reverse proxy works for the website, not the visitor. Requests to the site's address land on the reverse proxy first, and it decides what to do with them: serve a cached copy, pass the request to one of many application servers, terminate TLS, or refuse it. Nginx, HAProxy and Envoy are common self-hosted reverse proxies; CDNs such as Cloudflare, Akamai and Fastly run them at global scale.

Why scrapers meet them

You never buy a reverse proxy to scrape, but you talk to one on almost every large site. The forward proxy you pay for connects to the site's reverse proxy, and that is frequently where bot detection runs: the CDN reads your IP, TLS fingerprint and headers before the website's own servers see anything. A block page with a CDN's branding came from its reverse proxy, which is why Cloudflare errors look the same across thousands of unrelated sites.

How to tell what is in front of a site

Response headers usually give it away:

curl -sI https://example.com/ | grep -iE '^(server|via|cf-ray|x-cache)'

A server: cloudflare or cf-ray header points to Cloudflare; x-cache and a Via header often point to a caching layer or CDN.

Common confusion

Both kinds are "proxies", which causes mix-ups in search results and documentation. The test is whose configuration it lives in. If you set it in your client, it is a forward proxy. If the site operator set it up and you did nothing, it is a reverse proxy. What is a proxy server compares forward, reverse and transparent proxies side by side.