A Crawlee proxy is set with new ProxyConfiguration({ proxyUrls: [...] }), passed to any crawler as proxyConfiguration. Crawlee rotates the list round-robin and pins each session to one proxy, so a session's cookies always leave from the same IP. The same object works for CheerioCrawler, PlaywrightCrawler and the rest of the family.

const proxyConfiguration = new ProxyConfiguration({
    proxyUrls: ['http://USERNAME:PASSWORD@HOST1:PORT', 'http://USERNAME:PASSWORD@HOST2:PORT'],
});

Every snippet on this page ran on 2026-09-30 with Crawlee 3.18.2 on Node.js 22.23, Playwright 1.63 with its bundled Chromium, and Crawlee for Python 1.10.3, against two local authenticating HTTP proxies and a SOCKS5 proxy that enforces username and password.

Before you start: copy your proxy details

  1. Open your order in the dashboard at https://app.proxyhive.io.
  2. For each IP, copy HOST, PORT (the HTTP port), USERNAME and PASSWORD. Every ISP or datacenter IP is its own endpoint, so a list of IPs becomes a list of proxy URLs.
  3. Put them in one variable:
export PROXY_URLS="http://USERNAME:PASSWORD@HOST1:PORT,http://USERNAME:PASSWORD@HOST2:PORT"

A crawl spreads well over a handful of datacenter proxies; switch to ISP for targets that treat hosting ranges with suspicion.

Crawlee proxy setup with CheerioCrawler

import { CheerioCrawler, ProxyConfiguration } from 'crawlee';

const proxyConfiguration = new ProxyConfiguration({
    proxyUrls: process.env.PROXY_URLS.split(','),
});

const crawler = new CheerioCrawler({
    proxyConfiguration,
    maxConcurrency: 4,
    async requestHandler({ request, $, proxyInfo }) {
        console.log(request.url, proxyInfo.hostname, proxyInfo.port, $('title').text().trim());
    },
    failedRequestHandler({ request }, error) {
        console.log('failed', request.url, error.message);
    },
});

await crawler.run([
    'https://books.toscrape.com/',
    'https://books.toscrape.com/catalogue/page-2.html',
]);

proxyInfo tells the handler which proxy served the request, which is the first thing to log when one IP starts failing. Keep credentials in the URL; Crawlee sends them as Proxy-Authorization and tunnels HTTPS with CONNECT.

How Crawlee proxy rotation works

ProxyConfiguration.newUrl() is what the crawlers call. We called it directly to see the rules:

const config = new ProxyConfiguration({ proxyUrls: process.env.PROXY_URLS.split(',') });
console.log(await config.newUrl('a'), await config.newUrl('a'), await config.newUrl('b'));
console.log(await config.newUrl(), await config.newUrl());

With a session ID, the same ID got the same proxy every time (a, a on the first, b on the second). Without one, calls alternated. The Crawlee proxy management guide documents two more options: null in proxyUrls means "no proxy" for that slot, and tieredProxyUrls escalates to a pricier tier when a site starts blocking.

Session pool pairing

The session pool is on by default. Each session gets its own cookie jar and, through newUrl(session.id), its own proxy:

const crawler = new CheerioCrawler({
    proxyConfiguration,
    useSessionPool: true,
    persistCookiesPerSession: true,
    sessionPoolOptions: { maxPoolSize: 2 },
    async requestHandler({ session, proxyInfo }) {
        console.log(session.id, proxyInfo.port);
    },
});

Across ten requests with a pool of two, each session stayed on one proxy port. That pairing is what makes a crawl look like a few steady visitors rather than one visitor hopping between addresses. When a session trips a block, Crawlee retires it and the next session brings a different IP. Rotating vs static proxies covers when that behaviour helps.

Pick a proxy per target with newUrlFunction

Round-robin treats every URL alike. When one site should always leave from a specific IP, say because you allowlisted it there or a login lives on it, replace proxyUrls with a function:

const [primary, fallback] = process.env.PROXY_URLS.split(',');
const proxyConfiguration = new ProxyConfiguration({
    newUrlFunction: (sessionId, { request } = {}) =>
        request?.url.includes('books.toscrape.com') ? primary : fallback,
});

In our crawl of two sites, books.toscrape.com went through the first proxy and httpbin.org through the second. The function receives the session ID too, so you can build your own pinning on top of it.

PlaywrightCrawler proxy

The same configuration drives a browser:

import { PlaywrightCrawler, ProxyConfiguration } from 'crawlee';

const proxyConfiguration = new ProxyConfiguration({
    proxyUrls: process.env.PROXY_URLS.split(','),
});

const crawler = new PlaywrightCrawler({
    proxyConfiguration,
    maxConcurrency: 2,
    navigationTimeoutSecs: 45,
    async requestHandler({ request, page, proxyInfo }) {
        console.log(request.url, proxyInfo.port, (await page.locator('body').innerText()).trim());
    },
});

await crawler.run([
    { url: 'https://api.ipify.org?format=json&n=1' },
    { url: 'https://api.ipify.org?format=json&n=2' },
]);

Crawlee starts a small local forwarder for each proxy that carries credentials and points the browser at it. Two consequences we confirmed: authentication needs no extra code, and a socks5://USERNAME:PASSWORD@HOST:SOCKS5_PORT URL worked here, which bare Playwright refuses.

Crawlee for Python proxies

Crawlee for Python mirrors the API with snake_case names:

import asyncio
import os

from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
from crawlee.proxy_configuration import ProxyConfiguration


async def main() -> None:
    proxy_configuration = ProxyConfiguration(proxy_urls=os.environ["PROXY_URLS"].split(","))
    crawler = BeautifulSoupCrawler(proxy_configuration=proxy_configuration)

    @crawler.router.default_handler
    async def handler(context: BeautifulSoupCrawlingContext) -> None:
        title = context.soup.title.get_text(strip=True) if context.soup.title else ""
        print(context.request.url, context.proxy_info.port, title)

    await crawler.run(["https://books.toscrape.com/", "https://books.toscrape.com/catalogue/page-2.html"])


asyncio.run(main())

Install it with pip install "crawlee[beautifulsoup]". Version 1.10.3 validates proxy URLs as http or https and raised a ValidationError on socks5://.

Common Crawlee proxy errors

MessageCauseFix
Proxy URL protocol "socks5:" is not supportedSOCKS5 in an HTTP crawlerUse the HTTP port
Detected a session error, rotating session..., repeatedEvery session failing, often a 407 from wrong credentialsRe-copy the credentials; check proxyInfo
URL scheme should be 'http' or 'https' (Python)SOCKS5 URL in Crawlee for PythonUse the HTTP port
Navigation timeouts in PlaywrightCrawlerSlow first load through a new tunnelRaise navigationTimeoutSecs

With wrong credentials, Crawlee kept rotating sessions before failing the request, so a bad password looks like a flaky site in the logs. Log proxyInfo in failedRequestHandler and check the proxy first. Target-side codes are in proxy error codes.

Next steps

  • LLM-ready markdown from a browser crawl: Crawl4AI.
  • Python and a mature middleware stack: Scrapy.
  • Connection reference: the docs.