"Pool" is used for two things. A provider's pool is the whole network of exits behind a rotating gateway. Your own pool is the list of proxies your scraper chooses from. Both are judged the same way: how many usable, distinct addresses you get for your job.
Why headline sizes mislead
A provider's total counts addresses across every country, including ones that are offline right now. What matters is how many distinct exits you get in the location you need, over the time you run, and how spread across networks and subnets they are. Many addresses in one /24 behave like far fewer, because sites block by range.
Measure it yourself
Send a few hundred requests through the gateway and count what comes back:
import ipaddress
import requests
proxy = "http://USERNAME:PASSWORD@HOST:PORT"
ips = []
for _ in range(200):
r = requests.get("https://api.ipify.org?format=json",
proxies={"http": proxy, "https": proxy}, timeout=30)
ips.append(r.json()["ip"])
subnets = {ipaddress.ip_network(f"{ip}/24", strict=False) for ip in ips if ":" not in ip}
print(len(set(ips)), "distinct IPs,", len(subnets), "distinct /24s")
Repeat at a different time of day. How to test a proxy provider adds location and network checks.
Your own pool
With static proxies you manage the pool: rotate across it, bench an address that starts failing, and bring it back later. How to rotate proxies in Python shows a health-checked pool. A bigger pool is not better if the extra addresses are blocked on your target; track success rate per address and prune.