A Crawl4AI proxy is a ProxyConfig with server, username and password, passed to CrawlerRunConfig(proxy_config=...). Proxies live on the run config, so each crawl call can use a different one. For rotation, hand a list of ProxyConfig objects to RoundRobinProxyStrategy and pass that as proxy_rotation_strategy.
run_config = CrawlerRunConfig(
proxy_config=ProxyConfig(server="http://HOST:PORT", username="USERNAME", password="PASSWORD"),
)
Every snippet on this page ran on 2026-09-30 with Crawl4AI 0.9.4 on Python 3.12, driving the Chromium build for Playwright 1.63, against two local authenticating HTTP proxies and a SOCKS5 proxy that enforces username and password.
Before you start: copy your proxy details
- Open your order in the dashboard at https://app.proxyhive.io.
- Copy HOST, PORT (the HTTP port), USERNAME and PASSWORD for each IP. Each ISP or datacenter IP is its own endpoint.
- Install Crawl4AI and its browser, then export the details:
pip install crawl4ai
crawl4ai-setup
export PROXY_SERVER="http://HOST:PORT" PROXY_USER=USERNAME PROXY_PASS=PASSWORD
Browser crawls that log in or hold state suit static ISP proxies: the address stays the same for the whole term.
Crawl4AI proxy setup: ProxyConfig on the run config
This fetches a page through the proxy and prints its markdown. No LLM, no API key.
import asyncio
import os
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig
async def main():
proxy = ProxyConfig(
server=os.environ["PROXY_SERVER"],
username=os.environ["PROXY_USER"],
password=os.environ["PROXY_PASS"],
)
run_config = CrawlerRunConfig(proxy_config=proxy, cache_mode=CacheMode.BYPASS, page_timeout=30_000)
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
result = await crawler.arun("https://books.toscrape.com/", config=run_config)
print(result.success, result.status_code)
print(result.markdown.raw_markdown[:500])
asyncio.run(main())
Our run returned True 200 and the page as markdown, starting with [Books to Scrape](https://books.toscrape.com/index.html) We love being scraped!. The proxy log showed the CONNECT to books.toscrape.com. cache_mode=CacheMode.BYPASS matters while testing: Crawl4AI caches results, and a cached page never touches the proxy.
The deprecated BrowserConfig proxy argument
Older examples pass BrowserConfig(proxy="http://USERNAME:PASSWORD@HOST:PORT"). On 0.9.4 it still routes traffic, with a UserWarning that the parameter "is deprecated and will be removed in a future release". The Crawl4AI proxy docs point to CrawlerRunConfig.proxy_config.
Build ProxyConfig with keyword arguments rather than ProxyConfig.from_string() if a password may contain @. We fed from_string the URL http://user:p@ss@HOST:PORT and got the password p and a server of http://ss@HOST:PORT.
LLM-free markdown: fit_markdown with a pruning filter
raw_markdown is the whole page. To drop navigation and boilerplate before the text reaches an LLM or a vector store, add a content filter. It runs locally:
from crawl4ai import DefaultMarkdownGenerator, PruningContentFilterLXML
run_config = CrawlerRunConfig(
proxy_config=proxy,
cache_mode=CacheMode.BYPASS,
markdown_generator=DefaultMarkdownGenerator(content_filter=PruningContentFilterLXML(threshold=0.48)),
)
result = await crawler.arun("https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html", config=run_config)
print(len(result.markdown.raw_markdown), len(result.markdown.fit_markdown))
On that product page fit_markdown came out at 1,527 characters against 1,933 raw: the breadcrumb links and the cover image went, and the title, price, stock line and description stayed. On pages with heavy navigation the saving is larger. Fewer tokens in means a cheaper bill wherever the text goes next, and the proxy traffic is the same either way.
Use PruningContentFilterLXML, not the older PruningContentFilter: on 0.9.4 the older class emits a DeprecationWarning that names the lxml version as its replacement, with identical output.
Crawl4AI proxy rotation with RoundRobinProxyStrategy
from crawl4ai.proxy_strategy import RoundRobinProxyStrategy
proxies = ProxyConfig.from_env()
strategy = RoundRobinProxyStrategy(proxies)
run_config = CrawlerRunConfig(proxy_rotation_strategy=strategy, cache_mode=CacheMode.BYPASS)
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
results = await crawler.arun_many([f"https://httpbin.org/anything/{i}" for i in range(4)], config=run_config)
for r in results:
print(r.url, r.success, r.status_code)
ProxyConfig.from_env() reads the PROXIES variable as comma-separated HOST:PORT:USERNAME:PASSWORD entries, which matches one of the copy formats the dashboard offers:
export PROXIES="HOST1:PORT:USERNAME:PASSWORD,HOST2:PORT:USERNAME:PASSWORD"
All four requests succeeded and both proxy logs showed traffic. To keep one IP for a group of requests, such as a login and the pages after it, add proxy_session_id="account-1" to the run config: the strategy binds that ID to one proxy until it is released. Rotating vs static proxies explains when each shape fits.
Crawl4AI SOCKS5 proxy: no authentication
Crawl4AI drives Chromium through Playwright, so it inherits Playwright's limit. A ProxyConfig(server="socks5://HOST:SOCKS5_PORT", username=..., password=...) failed with Browser.new_context: Browser does not support socks5 proxy authentication. Use the HTTP port with credentials, or allowlist your IP on the order and use SOCKS5 without them. The Playwright guide has the details.
Common Crawl4AI proxy errors
| Symptom | Cause | Fix |
|---|---|---|
result.success is False with Page.goto: Timeout 60000ms exceeded | Wrong credentials: our run with a bad password timed out instead of reporting 407 | Re-copy the credentials; set page_timeout lower while testing |
Browser does not support socks5 proxy authentication | SOCKS5 with username and password | HTTP port, or allowlist and no credentials |
UserWarning: The 'proxy' parameter is deprecated | BrowserConfig(proxy=...) | CrawlerRunConfig(proxy_config=...) |
| Requests never reach the proxy | A cached result was served | cache_mode=CacheMode.BYPASS |
Always check result.success and result.error_message: Crawl4AI returns failures as results instead of raising. Target-side codes are covered in proxy error codes.
Next steps
- A JavaScript crawler with sessions pinned to proxies: Crawlee.
- Plain HTTP scraping without a browser: web scraping with proxies in Python.
- Connection reference: the docs.