Web scraping with proxies in Python comes down to four things: pass a proxy URL such as http://USERNAME:PASSWORD@HOST:PORT to requests, set a timeout on every call, retry transient failures with urllib3's Retry, and parse the HTML with BeautifulSoup. This tutorial builds that into a complete scraper, step by step, that collects all 1,000 books from the practice site books.toscrape.com across 50 pages, using four threads and a pool of static IPs, and writes them to CSV.

Every script on this page was run end to end on Python 3.12 with requests 2.34, urllib3 2.8 and beautifulsoup4 4.15, through a local authenticating proxy. The practice target exists for exactly this purpose, so you can hammer your code against it without bothering anyone.

What you need

  • Python 3.10 or newer (current requests and urllib3 releases require it).
  • Three packages: pip install requests beautifulsoup4 (urllib3 comes with requests).
  • One or more proxies. This tutorial uses static datacenter or ISP IPs, where each IP is its own endpoint with a host, port, username and password. You can buy a single IP, at $3.20/IP a month for one datacenter IP, and the rate per IP falls as you add more.

Open your order in the ProxyHive dashboard and copy the host, port, username and password. Keep them out of your code:

export PROXY_HOST="HOST"
export PROXY_PORT="PORT"
export PROXY_USER="USERNAME"
export PROXY_PASS="PASSWORD"

If your password contains @, : or /, URL-encode it with urllib.parse.quote(password, safe="") before it goes into a URL. Our Python requests setup guide covers SOCKS5 and the other connection options in more detail.

Step 1: send a request through a proxy in Python

requests takes a proxies dictionary that maps the target URL's scheme to the proxy to use. Note that both keys point at an http:// proxy URL. That is correct: for an HTTPS target, requests asks the proxy to open a tunnel with CONNECT, and TLS to the website runs inside it.

import os

import requests

proxy = (
    f"http://{os.environ['PROXY_USER']}:{os.environ['PROXY_PASS']}"
    f"@{os.environ['PROXY_HOST']}:{os.environ['PROXY_PORT']}"
)
proxies = {"http": proxy, "https": proxy}

resp = requests.get("https://httpbin.org/ip", proxies=proxies, timeout=(5, 20))
print(resp.json())

resp = requests.get("https://books.toscrape.com/", proxies=proxies, timeout=(5, 20))
print(resp.status_code, len(resp.content), "bytes")

The first request prints the IP the website sees. It should be the proxy's IP, not yours. If it prints your own address, the proxy is not being applied; if it raises ProxyError with 407 in the message, the credentials are wrong. Proxy error codes explains every failure you are likely to meet here.

The timeout=(5, 20) tuple is a 5-second connect timeout and a 20-second read timeout. Without it, requests waits forever on a stalled connection, and a scraper with no timeout eventually hangs on one bad page.

Step 2: parse the page with BeautifulSoup

Each book on a listing page sits in an article.product_pod element with the title in the link's title attribute, the price in p.price_color and the rating as a class name on p.star-rating.

import os

import requests
from bs4 import BeautifulSoup

proxy = (
    f"http://{os.environ['PROXY_USER']}:{os.environ['PROXY_PASS']}"
    f"@{os.environ['PROXY_HOST']}:{os.environ['PROXY_PORT']}"
)
proxies = {"http": proxy, "https": proxy}

resp = requests.get(
    "https://books.toscrape.com/catalogue/page-1.html", proxies=proxies, timeout=(5, 20)
)
resp.raise_for_status()
soup = BeautifulSoup(resp.content, "html.parser")

for pod in soup.select("article.product_pod")[:3]:
    title = pod.select_one("h3 a")["title"]
    price = pod.select_one("p.price_color").get_text(strip=True)
    rating = pod.select_one("p.star-rating")["class"][1]
    print(f"{price:>8}  {rating:<5}  {title}")

print(soup.select_one("li.current").get_text(strip=True))
  £51.77  Three  A Light in the Attic
  £53.74  One    Tipping the Velvet
  £50.10  One    Soumission
Page 1 of 50

One detail matters here: we pass resp.content (bytes) to BeautifulSoup, not resp.text. This server sends Content-Type: text/html without a charset, so requests falls back to ISO-8859-1 and resp.text turns every pound sign into £. Given bytes, BeautifulSoup reads the page's own <meta charset> and decodes correctly. The BeautifulSoup documentation covers the selector API used above.

Step 3: reuse a Session and set default headers

Calling requests.get opens a new connection every time. A requests.Session keeps connections to the proxy alive, holds cookies between requests, and lets you set proxies and headers once:

session = requests.Session()
session.headers.update(HEADERS)
session.proxies = {"http": proxy_url, "https": proxy_url}

The headers we send are a current browser User-Agent and an Accept-Language. The default python-requests/2.x user agent announces a script, and some sites refuse it on sight. Headers are not a disguise on their own: the guide on how to avoid getting blocked while scraping explains what else a site looks at, from request rate to TLS fingerprints.

Step 4: retry failures with urllib3 Retry

Networks fail, and a scraper that treats every hiccup as fatal loses data. urllib3's Retry, mounted on the session through an HTTPAdapter, retries for you:

retry = Retry(
    total=4,
    connect=2,
    backoff_factor=1,
    status_forcelist=[429, 500, 502, 503, 504],
    allowed_methods=["GET", "HEAD"],
    respect_retry_after_header=True,
)
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))

What each setting does:

  • total=4: at most four retries per request, of any kind.
  • connect=2: at most two of those for connection failures.
  • backoff_factor=1: wait roughly 1, 2, then 4 seconds between attempts, so a struggling server gets room.
  • status_forcelist: retry these status codes. 429 and 503 are the target asking you to slow down; 502 and 504 are usually passing trouble on the path.
  • respect_retry_after_header=True: if the site says Retry-After: 30, wait 30 seconds.

When the retries run out on a status code, requests raises requests.exceptions.RetryError. What Retry will not fix is a 407 on an HTTPS target: that surfaces as a ProxyError straight away, because wrong credentials do not improve with time.

Step 5: detect blocks before they poison your data

A blocked scraper often does not fail loudly. Many sites answer a bot with 200 OK and a CAPTCHA page, and a parser that finds zero products on that page writes nothing and moves on. Over a long run, you end up with a CSV that looks complete and is not.

Check each page for what should be there, and for what should not:

class Blocked(Exception):
    pass


def looks_blocked(soup):
    text = soup.get_text(" ", strip=True).lower()
    if any(s in text for s in ("captcha", "access denied", "unusual traffic")):
        return True
    return soup.select_one("article.product_pod") is None

Then treat 403 and 429 as blocks, let raise_for_status() handle other errors, and only parse pages that pass:

def scrape_page(page):
    url = f"{BASE}page-{page}.html"
    resp = session_for_thread().get(url, timeout=TIMEOUT)
    if resp.status_code in (403, 429):
        raise Blocked(f"{url} answered {resp.status_code}")
    resp.raise_for_status()
    soup = BeautifulSoup(resp.content, "html.parser")
    if looks_blocked(soup):
        raise Blocked(f"{url} answered 200 without the product grid")
    return list(parse_rows(url, soup))

The expected element (article.product_pod) is the strongest test. Marker words vary by site, so add the phrases your target's block page uses once you have seen one.

Step 6: scrape concurrently with ThreadPoolExecutor and a pool of static IPs

Fetching 50 pages one after another spends most of its time waiting on the network. concurrent.futures.ThreadPoolExecutor runs several fetches at once, and a pool of static IPs spreads them so no single IP carries the whole load.

Read the pool from an environment variable, falling back to the single proxy from step 1:

export PROXY_LIST="http://USERNAME:PASSWORD@HOST1:PORT1,http://USERNAME:PASSWORD@HOST2:PORT2"

Each ISP or datacenter IP on your order is its own endpoint, so the list is one URL per IP. The dashboard can copy them in USER:PASS@HOST:PORT form, which only needs http:// in front.

requests does not promise that a Session is safe to share between threads, so give each thread its own, and assign each new session the next proxy from the pool:

_local = threading.local()
_proxies = cycle(proxy_urls())
_lock = threading.Lock()


def session_for_thread():
    if not hasattr(_local, "session"):
        with _lock:
            _local.session = make_session(next(_proxies))
    return _local.session

With four workers and two IPs, each IP carries two threads. That ratio, concurrency per IP, is the number that decides whether a strict site rate limits you, so keep it low and raise it only while responses stay fast and clean. This round-robin is the simplest form of rotation; rotating proxies in Python goes further, with health checks that bench a failing IP and per-domain affinity.

Step 7: save the results to CSV

csv.DictWriter writes one row per book. Open the file with newline="" so the csv module controls line endings, and encoding="utf-8" so titles and pound signs survive on every platform:

with open("books.csv", "w", newline="", encoding="utf-8") as fh:
    writer = csv.DictWriter(fh, fieldnames=["title", "price", "rating", "in_stock", "url"])
    writer.writeheader()
    writer.writerows(sorted(rows, key=lambda r: r["url"]))

Threads finish in any order, so sorting before writing makes two runs produce comparable files.

The complete script for web scraping with proxies in Python

Here is everything assembled into one file, scrape_books.py:

import csv
import os
import threading
from concurrent.futures import ThreadPoolExecutor, as_completed
from itertools import cycle
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

BASE = "https://books.toscrape.com/catalogue/"
PAGES = range(1, 51)
TIMEOUT = (5, 20)
HEADERS = {
    "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
    "(KHTML, like Gecko) Chrome/129.0 Safari/537.36",
    "Accept-Language": "en-GB,en;q=0.9",
}
RATINGS = {"One": 1, "Two": 2, "Three": 3, "Four": 4, "Five": 5}


class Blocked(Exception):
    pass


def proxy_urls():
    listed = os.environ.get("PROXY_LIST")
    if listed:
        return [u.strip() for u in listed.split(",") if u.strip()]
    env = os.environ
    return [f"http://{env['PROXY_USER']}:{env['PROXY_PASS']}@{env['PROXY_HOST']}:{env['PROXY_PORT']}"]


def make_session(proxy_url):
    retry = Retry(
        total=4,
        connect=2,
        backoff_factor=1,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET", "HEAD"],
        respect_retry_after_header=True,
    )
    session = requests.Session()
    session.mount("https://", HTTPAdapter(max_retries=retry))
    session.mount("http://", HTTPAdapter(max_retries=retry))
    session.headers.update(HEADERS)
    session.proxies = {"http": proxy_url, "https": proxy_url}
    return session


_local = threading.local()
_proxies = cycle(proxy_urls())
_lock = threading.Lock()


def session_for_thread():
    if not hasattr(_local, "session"):
        with _lock:
            _local.session = make_session(next(_proxies))
    return _local.session


def looks_blocked(soup):
    text = soup.get_text(" ", strip=True).lower()
    if any(s in text for s in ("captcha", "access denied", "unusual traffic")):
        return True
    return soup.select_one("article.product_pod") is None


def parse_rows(url, soup):
    for pod in soup.select("article.product_pod"):
        link = pod.select_one("h3 a")
        rating = pod.select_one("p.star-rating")["class"][1]
        yield {
            "title": link["title"],
            "price": pod.select_one("p.price_color").get_text(strip=True),
            "rating": RATINGS.get(rating),
            "in_stock": "In stock" in pod.select_one("p.availability").get_text(),
            "url": urljoin(url, link["href"]),
        }


def scrape_page(page):
    url = f"{BASE}page-{page}.html"
    resp = session_for_thread().get(url, timeout=TIMEOUT)
    if resp.status_code in (403, 429):
        raise Blocked(f"{url} answered {resp.status_code}")
    resp.raise_for_status()
    soup = BeautifulSoup(resp.content, "html.parser")
    if looks_blocked(soup):
        raise Blocked(f"{url} answered 200 without the product grid")
    return list(parse_rows(url, soup))


def main():
    rows, failed = [], []
    with ThreadPoolExecutor(max_workers=4) as pool:
        futures = {pool.submit(scrape_page, p): p for p in PAGES}
        for future in as_completed(futures):
            page = futures[future]
            try:
                rows.extend(future.result())
            except (requests.RequestException, Blocked) as exc:
                failed.append(page)
                print(f"page {page} failed: {exc}")

    with open("books.csv", "w", newline="", encoding="utf-8") as fh:
        writer = csv.DictWriter(fh, fieldnames=["title", "price", "rating", "in_stock", "url"])
        writer.writeheader()
        writer.writerows(sorted(rows, key=lambda r: r["url"]))

    print(f"saved {len(rows)} books, {len(failed)} pages failed: {sorted(failed)}")


if __name__ == "__main__":
    main()

Run it:

python scrape_books.py
saved 1000 books, 0 pages failed: []

We ran this through two local proxies and got all 1,000 rows. We also ran it with a deliberately wrong password: every page failed with Tunnel connection failed: 407 Proxy Authentication Required, and the script still finished and reported all 50 failures, which is the behaviour you want. A scraper should tell you what it missed, not crash on the first bad page or skip it silently.

Failed pages are listed so you can rerun only those. For a long crawl, write rows as you go (or checkpoint every few pages) so a crash halfway does not throw away the first half.

Scaling up: httpx async, Scrapy and headless browsers

Threads are fine for hundreds of pages. For tens of thousands, three routes are worth knowing.

Async with httpx

httpx offers an async client with a requests-like API. One client per proxy, a semaphore to cap concurrency, and asyncio.gather to run the pages:

import asyncio
import os

import httpx
from bs4 import BeautifulSoup

PROXY = os.environ["PROXY_URL"]


async def fetch_titles(client, limit, page):
    async with limit:
        resp = await client.get(f"https://books.toscrape.com/catalogue/page-{page}.html")
    resp.raise_for_status()
    soup = BeautifulSoup(resp.content, "html.parser")
    return [a["title"] for a in soup.select("article.product_pod h3 a")]


async def main():
    limit = asyncio.Semaphore(8)
    timeout = httpx.Timeout(20.0, connect=5.0)
    async with httpx.AsyncClient(proxy=PROXY, timeout=timeout) as client:
        pages = await asyncio.gather(*(fetch_titles(client, limit, p) for p in range(1, 51)))
    print(sum(len(p) for p in pages), "titles")


asyncio.run(main())

Set PROXY_URL to the full http://USERNAME:PASSWORD@HOST:PORT. In httpx 0.28 the argument is proxy=; the older proxies= argument was removed, so code copied from older posts fails with a TypeError.

Scrapy

Once you need scheduling, deduplication, per-domain throttling and pipelines, stop building them yourself. Scrapy has all of them, plus AutoThrottle. Our Scrapy proxy setup guide shows where the proxy goes.

Headless browsers

If the data only appears after JavaScript runs, requests and BeautifulSoup will see an empty shell. Check the browser's network tab first: the page often loads its data from a JSON endpoint you can call directly, which is faster and cheaper than rendering. If not, use Playwright or Selenium with the same proxy.

Which proxy type for web scraping in Python

The code above does not change with the proxy type; the target decides which one you need.

Proxy typeHow it behavesGood fit
Datacenter (static)Your own IPs on hosting networks, priced per IPPractice targets, APIs, sites that do not check the network type
ISP (static)Your own IPs registered to consumer carriers, priced per IPSites that refuse server ranges but reward a stable identity
Residential (rotating)A new exit IP from a large pool, priced per GBBroad sweeps over strict targets where volume per IP must stay low

Start with the cheapest type that works and measure: run 200 requests, count clean pages versus blocks, and only move up when the numbers say so. The comparison of residential, ISP and datacenter proxies and our web scraping use case page go deeper on the trade-offs.

Whatever you point this at, scrape public data at a pace the site can absorb, and check the target's terms. Our allowed-use policy is public and plain about what the network may and may not be used for. This is not legal advice.