CAPTCHA stands for Completely Automated Public Turing test to tell Computers and Humans Apart. The image grids are the familiar form, but the common ones today, such as Google reCAPTCHA, hCaptcha and Cloudflare Turnstile, often run invisibly: they score browser signals in the background and only show a puzzle when the score is poor.

Why you see more of them through proxies

A CAPTCHA is what a site shows when bot detection is unsure, rather than sure enough to block. The usual triggers:

  • IP reputation. Addresses from hosting networks, or ones that other customers have pushed hard, start from a lower score.
  • Request rate. Many requests from one IP in a short time, the classic rate limiting signal.
  • Client signals. A bare HTTP client cannot run the JavaScript check at all, and a headless browser with automation flags fails it.

How to detect one in a scraper

A CAPTCHA page often comes back as 200 OK or 403, so status codes alone miss it. Check the body for the markers each widget leaves, such as g-recaptcha, h-captcha or cf-turnstile, or check that the element you came for is present. Count challenge pages as failures when you measure a proxy's success rate.

What to do about them

Treat a rising CAPTCHA rate as feedback. Slow down, keep one IP per session instead of rotating mid-flow, make the client look like a real browser, and move strict targets to a cleaner IP type. How to avoid getting blocked covers each of these in order.