The standard Scrapy proxy setup is one meta key: set request.meta["proxy"] to http://USERNAME:PASSWORD@HOST:PORT and Scrapy's built-in HttpProxyMiddleware handles the authentication header and the CONNECT tunnel for HTTPS. To spread requests over several IPs, add a small downloader middleware that fills that key for you.
yield scrapy.Request(url, meta={"proxy": "http://USERNAME:PASSWORD@HOST:PORT"})
The spider and middleware on this page ran on 2026-09-29 with Scrapy 2.19 against two local authenticating proxies, including a run where one of them rejected the credentials.
Before you start: copy your proxy details
- Open your order in the dashboard at https://app.proxyhive.io.
- Copy HOST, PORT, USERNAME and PASSWORD for each IP. Every ISP or datacenter IP is its own endpoint, with its own HTTP, HTTPS and SOCKS5 ports on the order. Use the HTTP port with Scrapy.
- The dashboard can copy each line as
USER:PASS@HOST:PORT. Prefixhttp://and you have a proxy URL. Store the list in an environment variable, one URL per comma:
export PROXY_LIST="http://USER:PASS@HOST1:PORT1,http://USER:PASS@HOST2:PORT2"
Datacenter proxies are the cheap way to feed a crawler when the target accepts server traffic; move a job to static ISP proxies when it starts challenging you.
How Scrapy's HttpProxyMiddleware works
HttpProxyMiddleware is on by default (HTTPPROXY_ENABLED = True). For each request it:
- Uses
request.meta["proxy"]if set. - Otherwise falls back to
http_proxy/https_proxyfrom the environment, read once at startup. - Strips any
user:passfrom the URL and sends it asProxy-Authorization: Basic ....
That third step is why response.meta["proxy"] reads http://HOST:PORT afterwards, with no password. The Scrapy documentation covers the details, including HTTPPROXY_AUTH_ENCODING for non-ASCII credentials.
For HTTPS pages, Scrapy opens a CONNECT tunnel and TLS runs end to end with the site. The proxy URL scheme stays http://.
Scrapy proxy authentication: credentials in the URL or an IP allowlist
Credentials in the URL are the normal route. Percent-encode @, : and / in passwords with urllib.parse.quote(password, safe="").
On a crawler box with a fixed address, the cleaner option is allowlist authentication: switch the IP to it on the order in the dashboard, add the server's public IP, and use http://HOST:PORT with no credentials. Nothing secret sits in your settings or logs.
A rotating proxy middleware for Scrapy
Create middlewares.py in your project:
import itertools
class RotatingProxyMiddleware:
def __init__(self, proxies):
self.pool = itertools.cycle(proxies)
@classmethod
def from_crawler(cls, crawler):
proxies = crawler.settings.getlist("PROXY_LIST")
if not proxies:
raise ValueError("PROXY_LIST is empty")
return cls(proxies)
def process_request(self, request, spider=None):
if "proxy" not in request.meta or request.meta.get("retry_times"):
request.meta["proxy"] = next(self.pool)
Register it in settings.py below 750, so it runs before HttpProxyMiddleware reads the meta key. At 800 our run sent the proxy without its credentials and got a 407 on every request:
import os
PROXY_LIST = os.environ["PROXY_LIST"].split(",")
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.RotatingProxyMiddleware": 350,
}
Two details matter. A request that already names a proxy keeps it, so you can still pin one request to one IP by hand. And a retried request gets a fresh proxy: in our run, a request whose first IP answered 407 was retried on the next IP and succeeded. Without the retry_times check, RetryMiddleware would resend the copy to the same failing address. For smarter policies (random choice, cooldowns, removing a dead IP) see how to rotate proxies in Python.
A spider needs nothing proxy-specific. On Scrapy 2.13 and later the entry point is async def start(); our 2.19 run ignored a spider that defined only the older start_requests():
import scrapy
class IpSpider(scrapy.Spider):
name = "ip"
async def start(self):
for i in range(4):
yield scrapy.Request(f"https://httpbin.org/ip?n={i}", dont_filter=True)
def parse(self, response):
yield {"proxy": response.meta.get("proxy"), "body": response.text}
Scrapy proxy settings: concurrency, AutoThrottle and retries
CONCURRENT_REQUESTS = 16
CONCURRENT_REQUESTS_PER_DOMAIN = 4
DOWNLOAD_TIMEOUT = 30
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_TARGET_CONCURRENCY = 2.0
RETRY_ENABLED = True
RETRY_TIMES = 3
RETRY_HTTP_CODES = [429, 500, 502, 503, 504, 522, 524, 408]
CONCURRENT_REQUESTS_PER_DOMAINcounts per target domain, not per proxy. Four IPs do not quadruple the safe rate against one site; they spread it.DOWNLOAD_TIMEOUTdefaults to 180 seconds, which is long for a proxied request. Thirty keeps a dead connection from holding a slot.- AutoThrottle adjusts the delay per domain from observed latency, so a struggling target slows the crawl instead of collecting timeouts.
- The default retry list already includes 429. Rotating on retry, as above, gives the retry a different address.
Staying under a site's limits is part of avoiding blocks while scraping, and our allowed-use policy asks the same.
Verify the exit IP
Run the spider above against https://httpbin.org/ip and check each item's origin against the IPs on your order. We used httpbin rather than api.ipify.org here: ipify returned 400 Bad Request to Scrapy's default request even without a proxy, which would send you debugging the wrong layer.
SOCKS5 and Scrapy
The default handler speaks HTTP proxies only. Scrapy 2.19 includes an httpx-based download handler that can use SOCKS when the socks extra is installed, but the module's own docstring says it is not recommended for production. Use the HTTP port; HTTP vs SOCKS5 explains why that costs you nothing for web scraping.
Common Scrapy proxy errors
| Log line | Cause | Fix |
|---|---|---|
TunnelError: Could not open CONNECT tunnel with proxy HOST:PORT [{'status': 407, ...}] | Wrong credentials, or an unencoded special character | Re-copy; quote() the password |
TunnelError: ... [{'status': 502, 'reason': b'Bad Gateway'}] | The proxy could not reach the target (we reproduced it with a hostname that does not resolve) | Check the URL; retry on another IP |
DownloadConnectionRefusedError: Connection was refused by other side | Wrong host or port | Re-copy the HTTP port from the order |
407 on every request although the credentials are right | Rotation middleware registered at 750 or higher, so it sets the proxy after HttpProxyMiddleware ran and the credentials never become a header | Register it below 750 |
What 403, 429 and 503 from the target mean is covered in proxy error codes.
Next steps
- New to Python scraping with proxies? Web scraping with proxies in Python starts from plain requests, and the Python requests guide covers the client in depth.
- Connection reference: the docs.