A crawler's job is discovery. It starts from a few seed URLs, fetches each page, extracts the links, and adds new ones to a queue, called the frontier, until it runs out or hits a limit. Search engines run crawlers at enormous scale; most scraping projects run a small one to find the product or listing pages they want to extract.

Crawler vs scraper

The words overlap, but they describe different jobs:

CrawlerScraper
GoalFind URLsExtract fields from pages
OutputA list of pagesStructured data
Hard partsScope, deduplication, politenessParsing, rendering, blocks

Most real projects do both: crawl a category to find items, then scrape each item. Frameworks such as Scrapy and Crawlee handle the queue, deduplication and retries for you.

Crawling politely through proxies

Crawlers can generate a lot of requests quickly, so the rules matter more:

  • Read and follow robots.txt.
  • Limit concurrency and add a delay per host, not just in total.
  • Stay in scope. A crawler that follows every link can wander into calendars, faceted search and infinite pagination.
  • Identify yourself with a descriptive User-Agent where the site welcomes known crawlers.

Proxies spread a crawl across addresses, but they do not make an aggressive crawl acceptable. Our allowed-use policy treats crawling that degrades a site like a denial-of-service attack.

Common confusion

"Bot", "spider" and "crawler" are used loosely for any automated client. Strictly, only software that follows links to discover pages is a crawler; a script that fetches a fixed list of URLs is a scraper.