Concurrency is the number of requests in flight at once. Proxy vendors often call it "threads". It sets your throughput together with latency: at 20 requests in flight and half a second per request, you complete about 40 requests a second. Double either the concurrency or the speed and throughput roughly doubles, until something pushes back.
The two limits that push back
- The provider's. Many proxy plans cap concurrent connections per IP or per account, and a request over the cap is refused or queued. The limit is often in the plan details rather than the headline, so ask for it in writing before you size a job.
- The target's. Websites watch how many requests arrive from one IP at a time. High concurrency through a single static IP looks like one very busy client and meets rate limiting quickly; the same concurrency spread across a pool does not.
How to control it in code
Cap it explicitly rather than letting a loop open as many connections as it can. In Python's asyncio, a semaphore does it:
import asyncio
limit = asyncio.Semaphore(10)
async def fetch(session, url):
async with limit:
async with session.get(url, proxy="http://USERNAME:PASSWORD@HOST:PORT") as r:
return await r.text()
The aiohttp guide has the full pattern, and the Node.js equivalent with p-limit is in web scraping with Node.js.
Common confusion
More concurrency is not always faster. Past the point where the target starts returning 429 or challenge pages, extra connections only raise the failure rate. Start low, raise it while your success rate holds, and keep per-host concurrency modest even when the total is high.