A scraper does three things: fetch a page, pull the fields you want out of it, and store them. Fetching can be a plain HTTP request or a headless browser when the content is rendered by JavaScript. Parsing uses CSS selectors, XPath or the page's embedded JSON. Storage is anything from a CSV file to a database with price history.
Where proxies fit
At small scale you may not need one. Proxies become necessary when:
- Per-IP limits bite. A site caps how many pages one address may fetch, and your job needs more.
- Location changes the data. Prices, stock and search results differ by country or city.
- Your server's IP is distrusted. Many sites challenge traffic from cloud providers on sight.
A proxy only changes the IP. Headers, TLS fingerprint and request rate still decide whether you are blocked; bot detection covers the layers.
Before you write a scraper
Check the browser's network tab for a JSON endpoint the page itself calls; it is faster and more stable than parsing HTML. Check for an official API or data export. Read the site's terms and robots.txt. Whether scraping a given site is lawful depends on what you collect and where; this is not legal advice, and is web scraping legal walks through the main cases. Our allowed-use policy says what we permit.
Getting started
Web scraping with proxies in Python builds a working requests and BeautifulSoup scraper, and web scraping with Node.js does the same with undici and cheerio. For the proxy side of the job, see web scraping proxies.
Common confusion
Scraping and crawling are different jobs: a web crawler discovers pages, a scraper extracts data from them. Most projects need a little of both.