To rotate proxies in Python, keep your proxy URLs in a list and pick a different one for each request: itertools.cycle gives you round-robin, random.choice gives you random order, and a small pool class that tracks failures lets you bench a bad proxy until it cools down. If you use a rotating residential pool instead, the provider rotates exit IPs for you behind a single endpoint, and there is no list to manage at all.
This tutorial covers the first case, a list of static endpoints such as ISP or datacenter IPs, with five working patterns. Each script was run through two local authenticating proxies (plus a deliberately broken one) against books.toscrape.com and httpbin.org, using Python 3.12, requests 2.34 and httpx 0.28.
Two kinds of rotation: yours and the provider's
A static proxy is one IP that stays yours for the order's term. Buy ten of them and you have ten endpoints, each with its own host, port and credentials. Rotation is your code's job: which IP gets which request.
A rotating proxy pool is a single endpoint that assigns a different exit IP from the provider's pool per request or per session. Rotation happens on the provider's side, so your code sends every request to the same host and port.
Neither is better in general. Static IPs give you a stable identity, a known location and a price per IP; a rotating pool gives you breadth and a price per GB. Our guide to rotating vs static proxies goes through when each wins. The rest of this page is about managing a list of static endpoints well.
Load the proxy list from an environment variable
On ProxyHive, each ISP or datacenter IP on an order is its own endpoint. The dashboard lists them and can copy them as USER:PASS@HOST:PORT; add http:// in front of each and join them with commas:
export PROXY_LIST="http://USERNAME:PASSWORD@HOST1:PORT1,http://USERNAME:PASSWORD@HOST2:PORT2,http://USERNAME:PASSWORD@HOST3:PORT3"
Every script below starts by reading that variable, so no credentials live in the code:
PROXIES = [u.strip() for u in os.environ["PROXY_LIST"].split(",") if u.strip()]
Round-robin rotation with itertools.cycle
Round-robin hands out proxies in order and starts again at the end. Load spreads evenly, and it is easy to reason about: with ten proxies, each one carries every tenth request.
import os
import threading
from itertools import cycle
import requests
PROXIES = [u.strip() for u in os.environ["PROXY_LIST"].split(",") if u.strip()]
_rotation = cycle(PROXIES)
_lock = threading.Lock()
def next_proxy():
with _lock:
return next(_rotation)
def get(url, **kwargs):
proxy = next_proxy()
kwargs.setdefault("timeout", (5, 15))
return requests.get(url, proxies={"http": proxy, "https": proxy}, **kwargs), proxy
if __name__ == "__main__":
for _ in range(4):
resp, proxy = get("https://httpbin.org/ip")
print(proxy.rsplit("@", 1)[-1], resp.json()["origin"])
The lock matters once several threads share the cycle. Python does not document itertools.cycle as thread-safe, and a lock around next() costs nothing next to a network round trip.
The function returns the proxy alongside the response. Keep that habit: when something goes wrong, you want to know which IP it went wrong on.
Random rotation with random.choice
Random choice spreads load about as evenly as round-robin over many requests, without a shared position to lock:
import os
import random
import requests
PROXIES = [u.strip() for u in os.environ["PROXY_LIST"].split(",") if u.strip()]
def get(url, **kwargs):
proxy = random.choice(PROXIES)
kwargs.setdefault("timeout", (5, 15))
return requests.get(url, proxies={"http": proxy, "https": proxy}, **kwargs), proxy
if __name__ == "__main__":
for _ in range(4):
resp, proxy = get("https://httpbin.org/ip")
print(proxy.rsplit("@", 1)[-1], resp.status_code)
The catch is that short runs are lumpy: the same IP can come up three times in a row, which is exactly the burst a rate limiter notices. For a handful of IPs, round-robin is the better default. If some IPs should carry more traffic (say, faster ones), random.choices(PROXIES, weights=...) gives you weighted selection.
A health-checked proxy pool that benches failing proxies
Both patterns above keep sending traffic to a proxy that has stopped working, so every Nth request fails. A pool that counts failures and benches a proxy for a cooldown fixes that. Save this as proxy_pool.py:
import os
import threading
import time
from dataclasses import dataclass
import requests
PROXIES = [u.strip() for u in os.environ["PROXY_LIST"].split(",") if u.strip()]
class NoHealthyProxy(Exception):
pass
@dataclass
class ProxyState:
url: str
failures: int = 0
benched_until: float = 0.0
class ProxyPool:
def __init__(self, urls, max_failures=2, cooldown=30.0, max_cooldown=600.0):
self._states = {url: ProxyState(url) for url in urls}
self._order = list(self._states)
self._next = 0
self._lock = threading.Lock()
self.max_failures = max_failures
self.cooldown = cooldown
self.max_cooldown = max_cooldown
def is_available(self, url):
return self._states[url].benched_until <= time.monotonic()
def get(self):
with self._lock:
for _ in range(len(self._order)):
url = self._order[self._next % len(self._order)]
self._next += 1
if self.is_available(url):
return url
raise NoHealthyProxy("every proxy is cooling down")
def report(self, url, ok):
with self._lock:
state = self._states[url]
if ok:
state.failures = 0
return
state.failures += 1
if state.failures >= self.max_failures:
strikes = state.failures - self.max_failures
pause = min(self.cooldown * 2**strikes, self.max_cooldown)
state.benched_until = time.monotonic() + pause
def status(self):
now = time.monotonic()
return {
s.url.rsplit("@", 1)[-1]: max(0, round(s.benched_until - now))
for s in self._states.values()
}
PROXY_FAULTS = (requests.exceptions.ProxyError, requests.exceptions.ConnectTimeout)
BLOCK_CODES = {403, 407, 429}
def fetch(pool, url, attempts=4):
last_error = None
for _ in range(attempts):
proxy = pool.get()
try:
resp = requests.get(url, proxies={"http": proxy, "https": proxy}, timeout=(5, 15))
except PROXY_FAULTS as exc:
pool.report(proxy, ok=False)
last_error = exc
continue
if resp.status_code in BLOCK_CODES:
pool.report(proxy, ok=False)
last_error = RuntimeError(f"{resp.status_code} via {proxy.rsplit('@', 1)[-1]}")
continue
pool.report(proxy, ok=True)
return resp
raise RuntimeError(f"{url} failed after {attempts} attempts") from last_error
if __name__ == "__main__":
pool = ProxyPool(PROXIES)
for page in range(1, 7):
resp = fetch(pool, f"https://books.toscrape.com/catalogue/page-{page}.html")
print(page, resp.status_code, pool.status())
How it behaves:
get()walks the list round-robin and skips any proxy still on the bench.- After
max_failuresconsecutive failures, a proxy is benched for 30 seconds. Each further failure after it comes back doubles the pause, up to ten minutes. One success resets the count. fetch()retries a failed request on the next available proxy, so a single bad IP costs you a retry, not a lost page.- If every proxy is benched,
get()raisesNoHealthyProxy. That is your signal to stop and look, not to hammer harder.
We ran it with three proxies: one working, one pointing at a closed port, and one with a wrong password. By the third page, both broken proxies were benched for 30 seconds, and all six pages came back 200 through the working one.
What counts as a proxy failure
The classification is the part people get wrong. Bench a proxy for:
- Connection failures to the proxy. In requests, both a refused connection and a connect timeout to the proxy surface as
ProxyError. - 407, which on an HTTPS target also arrives as a
ProxyError(Tunnel connection failed: 407). - Block responses from the target: 403 and 429, and a 200 that turns out to be a CAPTCHA page if you check for those.
Do not bench a proxy for a 404 or a 500 from the target. Those are the website's problem, and benching a healthy proxy for them shrinks your pool for no reason. Proxy error codes explains how to tell a proxy's error from the target's in detail.
Check every proxy before a run
A pre-flight check catches a typo in the list before it turns into a stream of retries. This one asks an IP echo service through each proxy in parallel:
import os
from concurrent.futures import ThreadPoolExecutor
import requests
PROXIES = [u.strip() for u in os.environ["PROXY_LIST"].split(",") if u.strip()]
def exit_ip(proxy):
try:
resp = requests.get(
"https://api.ipify.org?format=json",
proxies={"http": proxy, "https": proxy},
timeout=(5, 10),
)
resp.raise_for_status()
return proxy, resp.json()["ip"]
except requests.RequestException as exc:
return proxy, f"FAILED: {type(exc).__name__}"
with ThreadPoolExecutor(max_workers=8) as pool:
results = list(pool.map(exit_ip, PROXIES))
for proxy, result in results:
print(f"{proxy.rsplit('@', 1)[-1]:<24} {result}")
healthy = [p for p, r in results if not r.startswith("FAILED")]
print(f"{len(healthy)}/{len(PROXIES)} proxies ready")
With static IPs, each line should print that proxy's own address. If two proxies print the same exit IP, you have a duplicate in the list.
Per-domain affinity: keep one IP per site
Rotating on every request is right for stateless pages. It is wrong when the site keeps state: a session cookie, a cart, a login to your own account. An identity whose IP changes on every click looks less like a person, not more. Per-domain affinity gives each domain one proxy and keeps it until that proxy is benched:
import threading
from urllib.parse import urlsplit
import requests
from proxy_pool import PROXIES, ProxyPool
class DomainAffinity:
def __init__(self, pool):
self.pool = pool
self._assigned = {}
self._lock = threading.Lock()
def proxy_for(self, url):
host = urlsplit(url).hostname
with self._lock:
proxy = self._assigned.get(host)
if proxy is None or not self.pool.is_available(proxy):
proxy = self.pool.get()
self._assigned[host] = proxy
return proxy
if __name__ == "__main__":
pool = ProxyPool(PROXIES)
affinity = DomainAffinity(pool)
session = requests.Session()
urls = [
"https://books.toscrape.com/",
"https://httpbin.org/ip",
"https://books.toscrape.com/catalogue/page-2.html",
"https://httpbin.org/headers",
]
for url in urls:
proxy = affinity.proxy_for(url)
try:
resp = session.get(url, proxies={"http": proxy, "https": proxy}, timeout=(5, 15))
except requests.exceptions.ProxyError:
pool.report(proxy, ok=False)
continue
pool.report(proxy, ok=resp.status_code not in (403, 429))
print(urlsplit(url).hostname, proxy.rsplit("@", 1)[-1], resp.status_code)
In our run, both books.toscrape.com requests went out through one proxy and both httpbin.org requests through the other. This is the static-list version of what providers call a sticky session. It also keeps the maths honest: each site sees one IP, so the rate you configure for that site is the rate that IP sends.
Affinity pairs well with static ISP proxies for work on accounts you own, where one fixed IP per account is the point. Our ISP proxies are sold one IP at a time for that reason.
Rotating proxies with asyncio and httpx
In httpx, the proxy belongs to the client, not the request. So the async pattern is one AsyncClient per proxy, rotated per request, with a semaphore to cap concurrency:
import asyncio
import os
from contextlib import AsyncExitStack
from itertools import count
import httpx
PROXIES = [u.strip() for u in os.environ["PROXY_LIST"].split(",") if u.strip()]
PROXY_FAULTS = (httpx.ProxyError, httpx.ConnectError, httpx.ConnectTimeout)
URLS = [f"https://books.toscrape.com/catalogue/page-{n}.html" for n in range(1, 21)]
async def main():
async with AsyncExitStack() as stack:
clients = [
await stack.enter_async_context(
httpx.AsyncClient(proxy=p, timeout=httpx.Timeout(15.0, connect=5.0))
)
for p in PROXIES
]
turn = count()
limit = asyncio.Semaphore(4 * len(clients))
async def fetch(url):
start = next(turn)
for offset in range(len(clients)):
client = clients[(start + offset) % len(clients)]
try:
async with limit:
resp = await client.get(url)
except PROXY_FAULTS:
continue
resp.raise_for_status()
return url, len(resp.content)
raise RuntimeError(f"{url}: every proxy failed")
results = await asyncio.gather(*(fetch(u) for u in URLS), return_exceptions=True)
ok = [r for r in results if not isinstance(r, Exception)]
errors = [r for r in results if isinstance(r, Exception)]
print(f"{len(ok)} pages fetched, {len(errors)} errors")
for err in errors:
print(type(err).__name__, err)
asyncio.run(main())
Two details come from testing. First, our initial version rotated with a shared cycle() and retried on "the next" proxy. With many tasks interleaving, a task could draw the same dead proxy on every attempt, and 2 or 3 of 20 pages failed. Giving each task its own starting offset and walking the list from there guarantees every retry uses a different proxy; after that change, 20 of 20 pages came back with one of two proxies dead. Second, the argument is proxy=: httpx 0.28 removed proxies=, so older examples fail with a TypeError.
The semaphore is four requests per proxy. That is concurrency per IP, the number a target's rate limiter sees, so tune it per target rather than per machine. Our guide on how to avoid getting blocked while scraping covers pacing, backoff with jitter and block detection.
Rotating residential pools: the provider rotates for you
Everything above assumes you hold the list. With a rotating residential pool you do not: the order gives you one host, port, username and password, and the provider assigns the exit IP. Your code looks like the single-proxy example from our web scraping tutorial, with no cycle, no pool and no bench.
How the exit changes (per request, or held for a while as a session) is controlled by the provider, and every provider does it differently. On ProxyHive, check the options on your residential order in the dashboard rather than copying a format from another vendor's docs. Two honest notes: residential is priced per GB ($5.50/GB for a single gigabyte, with a 1 GB free trial), and it has no location selection today. If you need IPs in a specific country or city, static ISP or datacenter IPs let you pick them, and this page is how you rotate them.
Which rotation strategy to use
Each pattern above earns its place for a different job. Most real scrapers combine two: a health-checked pool underneath, with round-robin or affinity deciding which healthy proxy goes next.
| Strategy | Best for | Weak spot |
|---|---|---|
Round-robin (itertools.cycle) | Even load over a small, reliable list | Keeps sending traffic to a dead proxy |
Random (random.choice) | Large lists; no shared state to lock | Short runs are lumpy, so one IP can burst |
| Health-checked pool | Any list you run for more than a few minutes | A few dozen lines of code to own |
| Per-domain affinity | Sessions, cookies, accounts you own | One busy domain loads one IP |
| Provider-side rotation | Broad sweeps over strict targets | Behaviour is set by the provider, not your code |
If you only take one thing from this page, take the health-checked pool. A rotation that cannot notice a broken proxy turns one bad IP into a steady trickle of failed requests, and that trickle is easy to miss in a long run.
How many proxies you need
Rotation divides your traffic by the number of IPs, so the size of the list follows from two figures: how many pages you need, and how many requests a single IP can send to that target before it gets rate limited or blocked. The second figure is the target's, and you can only measure it.
A worked example with illustrative numbers:
- You need 30,000 pages a day from one site.
- You test one IP at increasing rates and find the site stays clean at one request every 10 seconds, and starts answering 429 at one every 5. That is 360 requests per hour, or 8,640 a day, per IP at the safe rate.
- 30,000 divided by 8,640 is about 3.5, so you need four IPs.
- Add headroom for benched IPs and bad days: five or six.
Measure again whenever the target changes its defences, and keep the per-IP rate fixed as you grow: more pages means more IPs, not a faster rate per IP. Static IPs are sold one at a time, so you can match the list to the arithmetic instead of a vendor's minimum; one ISP IP is $3.20/IP a month, and ISP proxy pricing shows how the rate falls as the list grows.
Common mistakes when you rotate proxies in Python
- Rotating instead of slowing down. Ten IPs at an abusive rate get you ten blocked IPs. Rotation spreads a sane rate; it does not make an insane one sane.
- Rotating mid-session. A logged-in session that hops IPs is a classic fraud signal. Use affinity.
- Benching for the target's errors. A 404 is not a dead proxy.
- Sharing an iterator between threads without a lock. Use a lock, or give each thread its own proxy as the tutorial's thread-local sessions do.
- Hard-coding credentials. Read them from the environment, as every script here does, so rotating a password is a config change.