A headless browser is Chrome, Firefox or WebKit with the window switched off, controlled by a library such as Playwright, Puppeteer or Selenium. It downloads the page, runs its JavaScript and builds the DOM exactly as a visible browser would, so you can scrape content that a plain HTTP client never sees because it is rendered in the browser.

When you need one, and what it costs

Reach for a headless browser when the data only appears after JavaScript runs, or when a site's challenge requires a real browser. Check first whether the page fetches its data from a JSON endpoint you could call directly; that is faster and far cheaper.

The cost matters on proxies billed per gigabyte. A browser loads images, fonts, scripts and trackers, often many times the size of the HTML. Block what you do not need:

page.route("**/*", lambda route: route.abort()
           if route.request.resource_type in ("image", "media", "font")
           else route.continue_())

Setting a proxy

Proxies are set at launch or per context. In Playwright for Python:

browser = p.chromium.launch(proxy={
    "server": "http://HOST:PORT",
    "username": "USERNAME",
    "password": "PASSWORD",
})

Chromium cannot send a password to a SOCKS5 proxy, so use the HTTP port or an IP allowlist there. The Playwright and Puppeteer guides cover per-context proxies.

Common confusion

Headless does not mean invisible to the site. Automation leaves traces, such as navigator.webdriver and differences in the browser fingerprint, and anti-bot scripts look for them. A proxy changes the IP, not those traces. Playwright vs Puppeteer vs Selenium compares how each tool handles waiting, stealth and proxies.