Before a client can connect to example.com, someone has to turn the name into an IP address. If your machine does that lookup itself, your own resolver, usually your ISP's or your office's, sees every hostname you visit, even though the request then travels through the proxy.

Which setups leak

  • HTTP proxies do not, for the request itself. The client sends the hostname to the proxy, in the request line or the CONNECT line, and the proxy resolves it.
  • SOCKS5 can go either way. The protocol lets a client send either a resolved IP or the hostname. Many clients resolve locally by default, which is the leak.

In curl and Python requests, the scheme decides:

curl -x socks5://USERNAME:PASSWORD@HOST:PORT https://example.com/    # resolves locally
curl -x socks5h://USERNAME:PASSWORD@HOST:PORT https://example.com/   # proxy resolves

The h in socks5h means the hostname goes to the proxy. In Firefox the equivalent is the "Proxy DNS when using SOCKS v5" checkbox, covered in the Firefox guide.

Why it matters for scraping

Privacy is only half of it. Large sites run on CDNs that answer DNS with a server near the resolver asking. Resolve locally in London while your exit IP is in Tokyo, and you may connect to a London edge node through a Japanese IP, which can change the content you get and looks odd to the site. Letting the proxy resolve keeps the lookup and the exit in the same place.

Common confusion

Online "DNS leak tests" are built for VPN users and check the whole system. For a scraper, check the client you run: curl -v with a socks5:// proxy prints a local resolution step before it connects, and with socks5h:// it does not.