A crawler's job is discovery. It starts from a few seed URLs, fetches each page, extracts the links, and adds new ones to a queue, called the frontier, until it runs out or hits a limit. Search engines run crawlers at enormous scale; most scraping projects run a small one to find the product or listing pages they want to extract.
Crawler vs scraper
The words overlap, but they describe different jobs:
| Crawler | Scraper | |
|---|---|---|
| Goal | Find URLs | Extract fields from pages |
| Output | A list of pages | Structured data |
| Hard parts | Scope, deduplication, politeness | Parsing, rendering, blocks |
Most real projects do both: crawl a category to find items, then scrape each item. Frameworks such as Scrapy and Crawlee handle the queue, deduplication and retries for you.
Crawling politely through proxies
Crawlers can generate a lot of requests quickly, so the rules matter more:
- Read and follow robots.txt.
- Limit concurrency and add a delay per host, not just in total.
- Stay in scope. A crawler that follows every link can wander into calendars, faceted search and infinite pagination.
- Identify yourself with a descriptive User-Agent where the site welcomes known crawlers.
Proxies spread a crawl across addresses, but they do not make an aggressive crawl acceptable. Our allowed-use policy treats crawling that degrades a site like a denial-of-service attack.
Common confusion
"Bot", "spider" and "crawler" are used loosely for any automated client. Strictly, only software that follows links to discover pages is a crawler; a script that fetches a fixed list of URLs is a scraper.