Every busy site limits how fast one client can ask for pages, to protect its servers and to spot automation. The limit is a count over a window, such as 60 requests a minute, and the key it counts against can be your IP, your login, a cookie or an API key. Common implementations are fixed windows, sliding windows and token buckets, which allow short bursts but cap the average.

How a limit shows itself

The polite signal is 429 Too Many Requests, often with a Retry-After header saying how long to wait. Many sites are less polite: they return 403, serve a CAPTCHA, slow their responses, or return a normal-looking page with the data missing. Watch for all of these, not only the status code.

How to stay under it

  1. Pace per host, with a delay between requests and modest concurrency.
  2. Back off when told. Honour Retry-After; otherwise wait exponentially longer after each refusal, with some random jitter.
  3. Cache what you already fetched.
  4. Spread load across IPs where the limit is per IP and your request rate is reasonable for the site.

Our allowed-use policy asks you to respect rate limits and not degrade the sites you collect from; aggressive crawling that has the effect of a denial of service is not allowed. How to avoid getting blocked has working backoff code.

Common confusion

IP rotation spreads a per-IP limit; it does nothing against a per-account or per-session limit, which follows you across every address. And a site's published API limits are a contract you agreed to, not a puzzle to route around.