Is web scraping legal? Collecting data that any logged-out visitor can see is lawful in many cases in the US and the EU, and US courts have narrowed the main anti-hacking law so that it rarely reaches public pages. But "public" is not a free pass. Terms you agreed to, copyright and database rights, data protection law, and the harm your traffic causes can each make a particular scrape unlawful, and a scrape can be fine under one of those and not another.

This is not legal advice. It is a map of the questions and the primary sources, current as of September 2026, written by a proxy provider for people who build scrapers. Laws differ by country and facts decide cases. For anything that matters, ask a lawyer in your jurisdiction.

Each row is a separate legal question, and each has to come out right:

QuestionWhere it usually landsMain sources
Is the data behind a login or other gate?Public pages are far safer than gated onesCFAA, Van Buren, hiQ; UK Computer Misuse Act
Did you agree to terms that forbid it?Breach of contract is the most common riskhiQ (2022 district ruling), Meta v. Bright Data, Ryanair
Is the content protected by copyright or a database right?Facts are free; expression and EU databases may not beFeist; EU Database Directive; EU text and data mining rules
Is any of it personal data?Needs a lawful basis under GDPR; CCPA duties may applyGDPR Articles 6 and 14; EDPB guidelines; CCPA
Does your traffic harm the site?Load and disruption create liability of their ownTrespass to chattels cases

Public data vs data behind a login

The single most useful line in scraping law is the login page. Content that anyone can load without an account is treated very differently from content behind credentials, a paywall or other technical barrier.

Scraping behind a login means you almost certainly accepted terms to get the account, so contract law applies. It also brings computer misuse laws into play, because you are passing a gate the owner put up. Creating fake accounts to get past it, as hiQ's contractors did, adds a separate breach.

Public pages avoid most of that. They do not avoid data protection law (a public profile is still personal data) or copyright (a public article is still someone's work).

Terms of service and contract law

Contract is where most scraping disputes are won and lost, and the US cases show why.

hiQ Labs v. LinkedIn is the case everyone cites, usually for half of what it said. In 2022 the Ninth Circuit, reconsidering after the Supreme Court's Van Buren decision, again upheld an injunction letting hiQ keep accessing public LinkedIn profiles, finding a serious question whether the CFAA's "without authorization" concept applies where no authorization is generally required. Then, on 4 November 2022, the district court held that hiQ had breached LinkedIn's User Agreement, which "unambiguously prohibits hiQ's scraping and unauthorized use of the scraped data", and that hiQ also breached it through contractors who created fake profiles. The case ended in December 2022 with a stipulated judgment against hiQ and a permanent injunction.

Meta v. Bright Data went the other way on its facts. On 23 January 2024 the Northern District of California granted summary judgment to Bright Data, reading Meta's terms as governing users while logged in, so they did not forbid logged-off scraping of public Facebook and Instagram data.

Put together: terms bind you when you agreed to them, typically by making an account, and courts read them for what they cover. A browsewrap notice in a footer is weaker than a signup checkbox, but do not build a business on that distinction.

The EU adds a twist. In Ryanair v. PR Aviation (C-30/14, 15 January 2015) the Court of Justice held that the Database Directive does not apply to a database protected neither by copyright nor by the database right, so the Directive's user exceptions do not stop its owner from limiting its use by contract. Less protection in law can mean more freedom to restrict by terms.

Computer misuse laws: the CFAA after Van Buren and hiQ

The US Computer Fraud and Abuse Act makes it a crime, and a civil wrong, to access a computer "without authorization" or to "exceed authorized access". For years companies argued that breaking a website's terms did both.

Van Buren v. United States (Supreme Court, decided 3 June 2021, 6 to 3) narrowed the second phrase. A person exceeds authorized access when they obtain information from particular areas of a computer, such as files, folders or databases, that are off-limits to them. The Court called it a "gates-up-or-down inquiry": misusing information you were entitled to reach is not enough. In a footnote the Court left open whether those gates must be technological or can be set by contracts or policies.

The Ninth Circuit's 2022 hiQ decision applied that reasoning to public websites: where a site lets anyone in, it is hard to say a particular visitor is "without authorization". That is persuasive in the Ninth Circuit, not binding nationwide.

The limits of that protection come from Facebook v. Power Ventures (9th Cir. 2016). Power kept accessing Facebook, including data behind users' logins, after Facebook sent a cease-and-desist letter and blocked its IP addresses; Power switched IPs. The court held that access after the letter was without authorization. The lesson for anyone running proxies: a written "stop" plus a block you route around is the fact pattern that loses.

The UK Computer Misuse Act 1990, section 1, makes it an offence to cause a computer to perform a function intending to secure access that is unauthorised, knowing it is. There is no UK equivalent of the hiQ line of cases to lean on, so treat explicit refusals, blocks and login walls as meaning what they say.

Facts are not copyrightable. The US Supreme Court said so in Feist v. Rural (1991): "No one may copyright facts or ideas." Prices, stock levels, flight times and match scores are facts. Their selection and arrangement can be protected if original, and the text, photos and reviews around them usually are.

The EU adds a database right that has no US equivalent. Under Article 7 of the Database Directive (96/9/EC), a maker who made a "substantial investment in either the obtaining, verification or presentation" of a database can stop extraction or re-use of the whole or a substantial part. Article 7(5) also catches "repeated and systematic extraction" of insubstantial parts that conflicts with normal exploitation of the database. A scraper that copies a price comparison site's listings page by page is exactly what that paragraph was written for. The UK kept an equivalent right after Brexit.

Text and data mining has its own EU rules. Article 4 of the Digital Single Market Directive allows mining lawfully accessible works unless the rightholder has "expressly reserved" that use, "such as machine-readable means in the case of content made publicly available online". A robots.txt rule or a machine-readable opt-out can therefore carry legal weight in the EU when you mine content.

In practice: extracting facts to analyse them is lower risk than republishing what you collected, and republishing someone's text, images or a substantial part of their database is where copyright claims come from.

Personal data: GDPR and CCPA

If what you collect identifies people (names, profile handles, emails, photos, reviews signed with a name), data protection law applies, and it applies to public data too.

Under GDPR you need a lawful basis under Article 6 before you collect. For scraping that is almost always legitimate interest, Article 6(1)(f), which requires a real interest, necessity, and a balancing test your interest has to win. Article 14 requires telling people you collected their data, subject to narrow exceptions. GDPR also reaches controllers outside the EU who monitor the behaviour of people in the EU.

The European Data Protection Board's Guidelines 03/2026 on web scraping, adopted for public consultation on 7 July 2026, are written for scraping to train generative AI but the reasoning travels. They expect precise collection criteria, filters that keep out data you do not need, excluding sites that "clearly oppose the scraping of their content" through robots.txt, CAPTCHAs or similar, and they note that the absence of a robots.txt file "does not amount to consent".

Under the CCPA, "publicly available" information is excluded from personal information, but the exclusion is narrower than it sounds: government records, and information lawfully made available by the consumer or through widely distributed media. The law applies to businesses over its thresholds, so check whether you are one before assuming it does not.

Our allowed-use policy forbids collecting personal data without a lawful basis under GDPR, CCPA or equivalent law, and says plainly that public availability is not a lawful basis by itself.

Does robots.txt make scraping illegal?

Not on its own. RFC 9309, the 2022 standard for robots.txt, says it outright: "These rules are not a form of access authorization."

It still matters legally, in three ways:

  1. Notice. It is written evidence of what the owner objects to, which matters in contract, trespass and misuse disputes.
  2. EU copyright. A machine-readable reservation is one way to opt out of the text and data mining exception.
  3. Data protection. The EDPB treats a robots.txt objection as a reason to leave a site out of personal-data scraping.

Respecting it costs little. Our guide to scraping without getting blocked includes a robots.txt check that runs through your proxy with the rest of your traffic.

Rate, load and harm

Even lawful collection becomes a problem when it hurts the site. In the US, trespass to chattels requires harm: eBay won an injunction against Bidder's Edge in 2000 partly because its crawling consumed eBay's capacity, and the California Supreme Court in Intel v. Hamidi (2003) confirmed that without damage or impaired functioning there is no claim. Load is what turns "reading a public page" into "interfering with someone's servers".

Keep your footprint small: pace per site, cache, fetch only what changed, and stop when you see rate-limit responses. Our guide to Cloudflare errors explains what 1015 and similar codes are telling you. Our policy counts crawling aggressive enough to degrade a site as denial of service.

Where a proxy fits in

Using a proxy is legal in most countries and it does not change the answer to any question above. It changes where your requests come from, which is often what the job needs: seeing a German storefront from Germany, or spreading polite traffic over several IPs.

What a proxy should not be is a way around a clear "no". The Power Ventures court held it against the scraper that it switched IPs after being blocked and told to stop. If a site has refused you, the proxy is not the answer; asking, or using their API, is. Our web scraping use case page shows the jobs proxies are meant for, and our comparison of a proxy vs a VPN covers the legality of the tools themselves.

A practical checklist before you scrape

  1. Is there an API, a data feed or a licence? Use it. It removes most of the questions below.
  2. Is the data public? If it needs a login, read the terms you accepted; assume they bind you.
  3. Have you been told to stop, or blocked? Then stop and talk to the owner. Do not rotate around it.
  4. Do the terms prohibit it, and did you accept them? Treat a signup agreement as a contract.
  5. Is it facts, or someone's expression? Collect facts; do not republish text, images or large parts of a database.
  6. Is any of it personal data? Document a lawful basis, collect only what you need, and plan how you will meet transparency duties. Skip sites that object.
  7. What does robots.txt say? Honour it, and treat it as the owner's stated wishes.
  8. Is your rate harmless? Pace, cache, back off, and watch for signs of load.
  9. Does your provider's policy allow it? Ours is public, and you can read the allowed-use policy before you spend anything.
  10. Is the stake large? Ask a lawyer who practises where you and the site operate.

Most scraping that passes these ten is ordinary data collection. The disputes that make case law tend to fail two or three of them at once.