Every site can publish /robots.txt, a list of rules grouped by crawler name:
User-agent: *
Disallow: /checkout/
Disallow: /search
Allow: /search/help
User-agent: ExampleBot
Disallow: /
A crawler finds the group that matches its user agent, or *, and skips the disallowed paths. The format was a convention for decades before it became RFC 9309 in 2022. Crawl-delay, which many sites use to ask for a pause between requests, is not part of the RFC, and crawlers treat it differently.
Check it in code
Python's standard library parses the file for you:
from urllib import robotparser
rp = robotparser.RobotFileParser("https://example.com/robots.txt")
rp.read()
print(rp.can_fetch("MyScraper/1.0", "https://example.com/products/"))
print(rp.crawl_delay("MyScraper/1.0"))
Scrapy obeys robots.txt by default in new projects through its ROBOTSTXT_OBEY setting.
Why it matters with proxies
A proxy changes where your requests come from, not what the site asked of you. robots.txt is not access control, and nothing technical stops a client from ignoring it, but it is the site's published statement of what it wants crawled. Honouring it is part of collecting responsibly, and whether you did can matter if a dispute arises. How much weight it carries legally depends on the jurisdiction and the facts; this is not legal advice. Is web scraping legal walks through the cases, and our allowed-use policy sets out what we permit.
Common confusion
An allowed path is not permission to fetch it as fast as you like. Rate limits, the site's terms and the load you put on its servers still apply, whatever robots.txt says.