To avoid getting blocked while scraping, behave like a considerate visitor rather than a flood: pace requests per site and per IP, back off with random jitter when you see 429 or 503, send consistent browser-like headers over a matching TLS fingerprint, keep cookies within a session, respect robots.txt and the site's terms, detect soft blocks, and cache what you already have. Then choose the proxy type the target expects. Proxies spread load and change what network you appear from; they do not fix a scraper that behaves like a bot.

That last point is worth saying plainly, because it is the one proxy vendors skip. If your scraper sends 50 requests a second with a default Python fingerprint, ten proxies give you ten blocked IPs. Sort out behaviour first, then add IPs.

Why scrapers get blocked

Anti-bot systems score each request and each client on a handful of signal families. Knowing which one tripped tells you what to fix:

SignalWhat the site measuresWhat fixes it
Rate and volumeRequests per IP, per session and per minute; burstsPacing, backoff, concurrency caps per IP
IP reputation and network typeWhether the IP belongs to a hosting network, a consumer carrier or a known proxy range; its historyThe right proxy type for the target
Request fingerprintHeader set and order, TLS handshake, HTTP/2 settingsConsistent headers and an impersonating TLS client
BehaviourNo cookies, no assets, perfectly regular timing, deep links with no referrerSessions, jitter, sensible navigation
Explicit rulesrobots.txt, terms of service, login wallsRespect them, or do not scrape that target

The first three cause most blocks in practice. The sections below take them in order, with Python you can run. We ran every script here through a local authenticating proxy against books.toscrape.com and httpbin.org.

Rate limiting and backoff

The biggest single lever is how fast you send. A site's rate limiter mostly counts requests per IP per window, so the rule is to pace per host, whatever your total concurrency.

This throttle hands out time slots per host, with a little randomness so requests do not arrive on a metronome:

class HostThrottle:
    def __init__(self, min_interval=2.0):
        self.min_interval = min_interval
        self._next_slot = {}
        self._lock = threading.Lock()

    def wait(self, host):
        with self._lock:
            now = time.monotonic()
            slot = max(now, self._next_slot.get(host, now))
            self._next_slot[host] = slot + self.min_interval * random.uniform(0.8, 1.2)
        time.sleep(slot - now)

When the site pushes back with 429 or 503, back off. Two rules make backoff work:

  1. Honour Retry-After. It can be a number of seconds or an HTTP date, so handle both.
  2. Add jitter. If 20 workers all wait exactly 2, 4, 8 seconds, they retry in lockstep and hit the limiter together again. "Full jitter", a random wait between zero and the exponential cap, spreads them out; AWS's write-up on exponential backoff and jitter shows why it beats the alternatives.
def retry_after(resp):
    value = resp.headers.get("Retry-After")
    if not value:
        return None
    if value.isdigit():
        return float(value)
    try:
        return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
    except (TypeError, ValueError):
        return None


def backoff(attempt, base=1.0, cap=60.0):
    return random.uniform(0, min(cap, base * 2**attempt))

Retries with jitter: what to retry and what not to

  • Retry 429, 503, 502, 504 and timeouts, with backoff.
  • Do not retry immediately on 403 from the same IP. The site has decided; asking again at once confirms you are a bot. Slow down, and if it persists, look at your fingerprint or the IP type.
  • Never retry 404 or 410. The page is not there.
  • Cap attempts. Five is plenty. A request that fails five times with backoff is telling you something.

If you are unsure which side sent an error, our guide to proxy error codes shows how to tell a proxy's 407 or 502 from a target's 403 or 429.

Concurrency per IP

Total concurrency is your business. Concurrency per IP is the site's. Ten threads spread over ten IPs look like ten calm visitors; ten threads on one IP look like one very busy bot.

Rules of thumb:

  • Start with one or two concurrent requests per IP per site, and raise it only while responses stay fast and clean.
  • Watch latency as well as errors. Many sites slow down suspected bots before they block them, so a rising response time is an early warning.
  • Keep the ratio fixed as you scale: more work means more IPs, not more threads per IP.

The rotating proxies in Python tutorial shows how to spread requests across a list of IPs, bench the ones that fail, and cap concurrency per proxy in async code.

Realistic headers and TLS fingerprints

Headers are the obvious part. The default python-requests/2.x User-Agent announces a script, and some sites refuse it outright. Send a current browser User-Agent, an Accept-Language, and the Accept header a browser would send for that kind of resource.

The less obvious part is below HTTP. Before any header is sent, your client performs a TLS handshake, and the details it offers (cipher suites, extensions, their order, supported groups) form a fingerprint. Methods such as JA3 and JA4 hash it; HTTP/2 connections add another fingerprint from their settings frames. Python's requests uses OpenSSL's defaults, which look nothing like Chrome's handshake. So a request claiming Chrome/129 in the User-Agent over an OpenSSL handshake is a mismatch, and mismatches are what anti-bot systems score.

Two ways out:

  1. Impersonate a browser's TLS stack. curl_cffi wraps a build of curl that reproduces browser handshakes and HTTP/2 settings, behind a requests-like API:
import os

from curl_cffi import requests

PROXY = os.environ["PROXY_URL"]

resp = requests.get(
    "https://httpbin.org/headers",
    impersonate="chrome",
    proxies={"http": PROXY, "https": PROXY},
    timeout=20,
)
print(resp.status_code, resp.json()["headers"]["User-Agent"])

With impersonate="chrome", curl_cffi sets the matching User-Agent and headers for you (version 0.16 sent a Chrome 150 User-Agent in our run), so do not override them with a different browser's.

  1. Use a real browser. For JavaScript-heavy targets, Playwright or Puppeteer produce a real browser's fingerprint because they are one. Our Playwright proxy setup guide shows the configuration. Browsers cost far more CPU per page, so reach for them when the data needs rendering, not by default.

Whichever you choose, consistency beats randomness. Rotating the User-Agent on every request within one session, while the TLS fingerprint stays the same, is its own tell. Pick one realistic profile per session and keep it.

Cookies and sessions

A browser that never keeps cookies is unusual. Use a requests.Session (or your client's equivalent) per logical visitor so cookies set on the first page travel with the next ones, and connections to the proxy are reused.

Tie each session to one IP. A session cookie issued to one IP and presented from another is a common fraud signal, which is why the rotation tutorial includes per-domain affinity. With static ISP or datacenter IPs, affinity is free: the IP does not change unless you change it. With a rotating residential pool, how long an exit IP is held depends on the options on your order in the dashboard.

robots.txt and terms of service

robots.txt is a site telling automated clients what it would rather they skip, standardised in RFC 9309. Python ships a parser:

def allowed_by_robots(url):
    parts = urlsplit(url)
    root = f"{parts.scheme}://{parts.netloc}"
    if root not in _robots:
        parser = RobotFileParser()
        resp = session.get(f"{root}/robots.txt", timeout=10)
        parser.parse(resp.text.splitlines() if resp.status_code == 200 else [])
        _robots[root] = parser
    return _robots[root].can_fetch(USER_AGENT, url)

Fetching robots.txt through your own session, rather than with RobotFileParser.read(), means it goes through the same proxy and headers as everything else. parser.crawl_delay(USER_AGENT) returns a Crawl-delay when the site sets one; use it as your minimum interval.

Terms of service matter more than robots.txt. If a site forbids automated access, or the data sits behind a login you are not entitled to automate, no proxy changes that. Our allowed-use policy is public and plain: public data collection is allowed if you respect the target's rate limits and do not degrade its service; crawling aggressive enough to degrade a site counts as denial of service; and personal data needs a lawful basis, since being public is not one by itself. This is not legal advice.

Detect soft blocks: CAPTCHA pages and 200s that are not

The block that costs most is the one you do not notice. Many sites answer a suspected bot with 200 OK and a challenge page, a CAPTCHA, or a stripped-down page with no data. Your parser finds nothing, writes nothing, and the run reports success.

Check every response for what should be there, not only the status code:

BLOCK_MARKERS = ("captcha", "access denied", "unusual traffic", "are you a robot")


def soft_blocked(resp, must_contain):
    body = resp.text.lower()
    return must_contain not in body or any(m in body for m in BLOCK_MARKERS)

Signs of a soft block, roughly from most to least reliable:

  1. The element you expect (a product grid, a price, a results list) is missing.
  2. The body contains challenge or CAPTCHA wording.
  3. You were redirected to a path such as /captcha, /challenge or /blocked.
  4. The response is much smaller than a normal page from the same template.
  5. The same URL returns different content through different IPs.

Log soft blocks per proxy and per target. A rising rate on one IP means that IP is burnt for that target; a rising rate on every IP means your behaviour or fingerprint is the problem.

Choose the proxy type per target

With pacing and fingerprint in order, the IP type decides the rest. Match it to what the target checks:

TargetUsual first choiceWhy
Public APIs, practice sites, lenient sitesDatacenter (static)Cheapest per request; nothing checks the network type
Sites that refuse hosting ranges but reward a stable identityISP (static)Consumer-carrier registration with a fixed IP, and you choose country, region, city and carrier
Strict targets where volume per IP must stay lowResidential (rotating)Each request can leave from a different consumer IP
Accounts you ownISP (static)One fixed IP per account

Start with the cheapest type that works and measure it: send a few hundred requests at your planned rate and count clean pages against blocks. Move up only when the numbers say so. Our comparison of residential, ISP and datacenter proxies goes deeper. One limit to know: ProxyHive residential has no location selection today, so for country-exact results use ISP or datacenter IPs.

Cache what you already have

The request you do not send cannot be blocked. Two habits cut volume sharply on repeat crawls:

  1. Do not refetch what has not changed. Many servers send ETag or Last-Modified. Send them back as If-None-Match and If-Modified-Since, and an unchanged page returns 304 Not Modified with no body:
import json
import os
from pathlib import Path

import requests

PROXY = os.environ["PROXY_URL"]
CACHE = Path("http-cache.json")


def load():
    return json.loads(CACHE.read_text()) if CACHE.exists() else {}


def cached_get(session, url, cache):
    headers = {}
    entry = cache.get(url)
    if entry:
        if entry.get("etag"):
            headers["If-None-Match"] = entry["etag"]
        if entry.get("last_modified"):
            headers["If-Modified-Since"] = entry["last_modified"]
    resp = session.get(url, headers=headers, timeout=(5, 20))
    if resp.status_code == 304:
        return entry["body"], True
    resp.raise_for_status()
    cache[url] = {
        "etag": resp.headers.get("ETag"),
        "last_modified": resp.headers.get("Last-Modified"),
        "body": resp.content.decode("utf-8", "replace"),
    }
    return cache[url]["body"], False


if __name__ == "__main__":
    session = requests.Session()
    session.proxies = {"http": PROXY, "https": PROXY}
    cache = load()
    for _ in range(2):
        body, from_cache = cached_get(session, "https://books.toscrape.com/catalogue/page-1.html", cache)
        print(len(body), "from cache" if from_cache else "downloaded")
    CACHE.write_text(json.dumps(cache))

On books.toscrape.com, the second request came back as a 304 and was served from the cache. For production, the requests-cache library does the same with proper storage backends.

  1. Crawl smarter, not wider. Use sitemaps and listing pages to find what changed, and fetch detail pages only for those. On proxies priced per GB, every page you skip is also money you keep.

Putting it together in Python

Here is the whole polite fetcher in one file: per-host throttle, robots.txt check, backoff with jitter that honours Retry-After, and soft-block detection. Set PROXY_URL to http://USERNAME:PASSWORD@HOST:PORT from your order.

import os
import random
import threading
import time
from email.utils import parsedate_to_datetime
from urllib.parse import urlsplit
from urllib.robotparser import RobotFileParser

import requests

PROXY = os.environ["PROXY_URL"]
USER_AGENT = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0 Safari/537.36"
BLOCK_MARKERS = ("captcha", "access denied", "unusual traffic", "are you a robot")

session = requests.Session()
session.headers["User-Agent"] = USER_AGENT
session.proxies = {"http": PROXY, "https": PROXY}


class HostThrottle:
    def __init__(self, min_interval=2.0):
        self.min_interval = min_interval
        self._next_slot = {}
        self._lock = threading.Lock()

    def wait(self, host):
        with self._lock:
            now = time.monotonic()
            slot = max(now, self._next_slot.get(host, now))
            self._next_slot[host] = slot + self.min_interval * random.uniform(0.8, 1.2)
        time.sleep(slot - now)


def retry_after(resp):
    value = resp.headers.get("Retry-After")
    if not value:
        return None
    if value.isdigit():
        return float(value)
    try:
        return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
    except (TypeError, ValueError):
        return None


def backoff(attempt, base=1.0, cap=60.0):
    return random.uniform(0, min(cap, base * 2**attempt))


def soft_blocked(resp, must_contain):
    body = resp.text.lower()
    return must_contain not in body or any(m in body for m in BLOCK_MARKERS)


_robots = {}


def allowed_by_robots(url):
    parts = urlsplit(url)
    root = f"{parts.scheme}://{parts.netloc}"
    if root not in _robots:
        parser = RobotFileParser()
        resp = session.get(f"{root}/robots.txt", timeout=10)
        parser.parse(resp.text.splitlines() if resp.status_code == 200 else [])
        _robots[root] = parser
    return _robots[root].can_fetch(USER_AGENT, url)


throttle = HostThrottle(min_interval=1.0)


def polite_get(url, must_contain, attempts=5):
    if not allowed_by_robots(url):
        raise PermissionError(f"robots.txt disallows {url}")
    host = urlsplit(url).hostname
    for attempt in range(attempts):
        throttle.wait(host)
        resp = session.get(url, timeout=(5, 20))
        if resp.status_code in (429, 503):
            delay = retry_after(resp) or backoff(attempt)
        elif resp.status_code == 200 and soft_blocked(resp, must_contain):
            delay = backoff(attempt, base=5.0)
        else:
            resp.raise_for_status()
            return resp
        if attempt == attempts - 1:
            break
        print(f"{url}: {resp.status_code}, waiting {delay:.1f}s")
        time.sleep(delay)
    raise RuntimeError(f"gave up on {url} after {attempts} attempts")


if __name__ == "__main__":
    for page in (1, 2, 3):
        r = polite_get(f"https://books.toscrape.com/catalogue/page-{page}.html", "product_pod")
        print(page, r.status_code, len(r.content))
    try:
        polite_get("https://httpbin.org/deny", "anything")
    except PermissionError as exc:
        print(exc)
    try:
        polite_get("https://httpbin.org/status/429", "x", attempts=3)
    except RuntimeError as exc:
        print(exc)
    try:
        polite_get("https://httpbin.org/html", "product_pod", attempts=2)
    except RuntimeError as exc:
        print(exc)

The demo at the bottom exercises each path. In our run, the three book pages came back 200; httpbin.org/deny was refused before any request because that site's robots.txt disallows it; the permanent 429 was tried three times, with a random wait of under two seconds between attempts, and then given up; and a page missing the expected element was treated as a soft block. For a full scraper built on the same ideas, with threads, a pool of static IPs and CSV output, follow the Python web scraping tutorial.

How to avoid getting blocked while scraping: a checklist

  1. Is there an API or a data export? Use it instead.
  2. Do robots.txt and the terms allow this? If not, stop.
  3. Is the pace per host and per IP set, with jitter?
  4. Does backoff honour Retry-After and cap attempts?
  5. Do headers and TLS fingerprint tell the same story?
  6. Does each session keep its cookies and its IP?
  7. Does every response get checked for the expected content?
  8. Are unchanged pages served from cache?
  9. Is the proxy type chosen by measurement, not by habit?

Get those right and most blocks never happen. The ones that still do are telling you something about the target, and that is worth listening to.