To scrape Google search results today, drive a real browser: plain HTTP clients now get a page that asks for JavaScript instead of results. Open https://www.google.com/search?q=...&hl=en&gl=us in Chrome through Playwright with a proxy, click through the consent page if you get one, parse each organic result by anchoring on the h3 inside a link, and stop the moment Google serves its "unusual traffic" page. Keep it slow, keep raw HTML, and check the official API first.

This tutorial builds that scraper in Python, one function at a time, then assembles it. It is written for rank tracking: a few hundred keywords, a few pages each, on a schedule. It is not a recipe for pulling millions of queries.

Before you scrape Google: terms, robots.txt and the official API

Three facts to weigh before writing any code.

Google's terms. The Google Terms of Service (effective 30 July 2026) list among forbidden uses "using automated means to access content from any of our services in violation of the machine-readable instructions on our web pages", and Google's robots.txt says Disallow: /search. Scraping result pages goes against those terms. What that means for you depends on your jurisdiction, your contract with Google and what you do with the data. This is not legal advice.

Our policy. Search engine result pages are an allowed use of our network: rank tracking, SERP monitoring and keyword research in any locale we sell, as our allowed-use policy states. That covers what you may do with our IPs. It does not change your agreement with Google.

The sanctioned route. Google's Custom Search JSON API returns results as JSON with no parsing at all. When we checked it on 2026-09-30, the page said it is closed to new customers, and existing customers have until 1 January 2027 to move off it. It gives 100 queries a day free and charges $5 per 1,000 after that, up to 10,000 a day. Google points new users at Vertex AI Search, which covers up to 50 domains, not the whole web. If you have access to the JSON API, use it for as long as it lasts. Everyone else is left with third-party SERP APIs or a scraper of their own, and this page is about the second.

Why plain requests fails: SERP scraping in Python in 2026

The first thing most people try is requests with a browser User-Agent:

import os

import requests

proxy = os.environ["PROXY_URL"]
resp = requests.get(
    "https://www.google.com/search",
    params={"q": "web scraping proxies", "hl": "en", "gl": "us"},
    headers={
        "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
        "(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
        "Accept-Language": "en-US,en;q=0.9",
    },
    proxies={"http": proxy, "https": proxy},
    timeout=20,
)
print(resp.status_code, len(resp.content), "<h3" in resp.text, "enablejs" in resp.text)

On 2026-09-30 that returned 200 and a 92 KB page with no result headings at all. It contained a <noscript> redirect to /httpservice/retry/enablejs: Google has required JavaScript for search since January 2025. A status code of 200 tells you nothing here, which is the first lesson of SERP scraping: always check the page for what should be on it.

So the fetching half of this tutorial uses Playwright to drive Chrome. The parsing half is still BeautifulSoup, on the HTML the browser ends up with.

What you need

  1. Python 3.10+, then pip install playwright beautifulsoup4. We used Playwright 1.63 and beautifulsoup4 4.15.
  2. A browser. We drove installed Google Chrome with channel="chrome". To use Playwright's bundled Chromium instead, run playwright install chromium and drop the channel argument.
  3. A proxy in the country whose results you want. Copy the host, port, username and password from your order in the dashboard and export them as PROXY_HOST, PROXY_PORT, PROXY_USER and PROXY_PASS. Use the HTTP port: Chromium cannot send a username and password to a SOCKS5 proxy. The Playwright proxy guide covers per-context proxies if you run several IPs from one browser.

Step 1: build the Google search URL

Everything the scraper asks for goes in the query string. Five parameters matter:

ParameterWhat it doesExample
qThe search queryq=web+scraping+proxies
hlInterface language (labels, buttons, consent text)hl=en
glCountry to bias results towardsgl=us, gl=de
startOffset of the first result, in steps of 10start=10 for page 2
numResults per page, historically up to 100Do not rely on it

num used to be the rank tracker's favourite shortcut. Google stopped honouring num=100 in September 2025 (it told reporters the parameter was never officially supported), so plan for ten organic results per page and paginate with start.

def search_url(query, page):
    params = {"q": query, "hl": HL, "gl": GL}
    if page > 1:
        params["start"] = (page - 1) * 10
    return "https://www.google.com/search?" + urlencode(params)

urlencode takes care of spaces and symbols in the query. Keep hl fixed across runs, because the consent handler and your block detection both match English text.

Step 2: open the results page through a proxy with Playwright

Playwright takes the proxy at launch. Keep the credentials in username and password, not in the server URL:

proxy = {
    "server": f"http://{os.environ['PROXY_HOST']}:{os.environ['PROXY_PORT']}",
    "username": os.environ["PROXY_USER"],
    "password": os.environ["PROXY_PASS"],
}
browser = p.chromium.launch(channel="chrome", proxy=proxy)
context = browser.new_context(locale="en-US", viewport={"width": 1366, "height": 900})

Then load the page, deal with consent, and wait for a result heading. A missing heading is not an error at this point; the block check decides what it means.

def fetch_serp(page, url):
    page.goto(url, wait_until="domcontentloaded", timeout=45_000)
    accept_consent(page)
    try:
        page.wait_for_selector("a[href] h3", timeout=10_000)
    except Exception:
        pass
    html = page.content()
    check_page(page.url, html)
    return html

A browser with no Google cookies, arriving from an EU or UK IP, usually meets a cookie consent step before any results: either a redirect to consent.google.com or a dialog over the page. Until someone clicks, there is nothing to parse. The handler looks for the "Reject all" button by its accessible name, so it does not depend on class names:

def accept_consent(page):
    button = page.get_by_role("button", name="Reject all", exact=True).first
    if urlsplit(page.url).hostname.startswith("consent.") or button.is_visible():
        button.click()
        page.wait_for_url(re.compile(r"^https://www\.google\.[a-z.]+/search"), timeout=15_000)
        page.wait_for_load_state("domcontentloaded")

Two honest caveats. The button text follows hl, so with hl=de you would match German labels instead. And our test IP never received the consent page during this run, so we exercised this handler against a mock consent page served through Playwright's request routing, not against Google's live one. Clicking "Reject all" is the choice that stores the least about your scraper, and once clicked the context keeps the cookie, so reuse the same context for the rest of the run.

Step 3: parse organic results with selectors that survive class churn

Google's class names (MjjYud, yuRUbf, VwiC3b and friends) are generated, and they change without notice. Tutorials that select div.g or .yuRUbf a break every few months. What has stayed put for years is structure: each organic result is a link whose visible title is an h3 inside it. Anchor on that.

def parse_results(html):
    soup = BeautifulSoup(html, "html.parser")
    results, seen = [], set()
    for h3 in soup.select("a[href] h3"):
        link = h3.find_parent("a")
        href = link["href"]
        if not is_organic(href) or href in seen:
            continue
        seen.add(href)
        block = link
        while block.parent is not None and len(block.parent.select("a[href] h3")) == 1:
            block = block.parent
        title = h3.get_text(" ", strip=True)
        results.append({"title": title, "url": href, "snippet": snippet_for(block, title)})
    return results

How it works:

  1. a[href] h3 finds every result title. The parent a gives the URL.
  2. is_organic drops links back into Google (Maps, Images, "More results") by host.
  3. To find the snippet, the parser climbs from the link to the largest ancestor that still holds exactly one result title. That ancestor is the result's card, whatever its class is called this month.
  4. Inside the card, snippet_for takes the longest text block that contains neither the title nor the cite breadcrumb.
def is_organic(href):
    host = urlsplit(href).hostname or ""
    return href.startswith("http") and not host.endswith(("google.com", "googleusercontent.com"))


def snippet_for(block, title):
    best = ""
    for el in block.find_all(["div", "span"]):
        if el.find("h3") or el.find("cite"):
            continue
        text = el.get_text(" ", strip=True)
        if title not in text and len(text) > len(best):
            best = text
    return best

What we tested the parser on

Our own IP was flagged before we got a clean results page (more on that below), so we did not keep fetching. Instead we parsed a saved Google results page published as a test fixture by an open-source SERP parser, captured in July 2026, with five organic results. The parser returned all five titles and URLs in page order, with the right snippets:

Machine code | https://en.wikipedia.org/wiki/Machine_code | In computing, machine code is data encoded and structured to control a
What Is Machine Language? (Explained for Abso | https://medium.com/@TechiesSpot/what-is-machine-la | Machine language is the lowest-level programming language . It's made
What is Machine Language? Assembler Explained | https://www.lenovo.com/us/en/glossary/machine-lang | Machine language is the lowest level of programming language that dire

Treat that as a snapshot. Google's markup changes, and this parser will need attention one day. Because every run saves raw HTML (step 5), you can fix the parser and re-parse old pages without fetching them again.

Titles are the easy part; ads, "People also ask", video carousels and AI overviews are harder. Ads sit in their own block and usually link through Google's click tracking, so is_organic drops most of them. Anything you need beyond organic rank deserves its own parser and its own saved examples.

Step 4: detect CAPTCHA and "unusual traffic" pages

This is the part that protects your data. When Google decides an IP is automated, it redirects to /sorry/index with a reCAPTCHA form and a message about unusual traffic from your network. A scraper that only counts results writes "zero results" for every keyword from then on.

def check_page(url, html):
    if "/sorry/" in url or 'id="captcha-form"' in html:
        raise Blocked(f"CAPTCHA page at {url[:80]}")
    if "<h3" not in html:
        if "unusual traffic" in html:
            raise Blocked("unusual traffic page")
        if "/httpservice/retry/enablejs" in html:
            raise Blocked("Google served its enable-JavaScript page")

We saw both pages for real on 2026-09-30. The requests call above got the JavaScript wall. The first headless Chrome visit that got through, from the test machine's own IP, landed on /sorry/index, and a second browser a minute later on the same IP got the same page, served with HTTP 429. check_page catches both saved copies.

When it fires, stop that IP. Do not solve the CAPTCHA automatically and do not retry in a loop: every request from a flagged IP keeps the flag fresh. Pause the IP, lower the rate, and come back later. Our guide on how to avoid getting blocked while scraping covers the pacing side, and TLS fingerprinting explains why a real browser gets further than an HTTP client wearing a browser's User-Agent.

Step 5: paginate, pace and store the results

The run loop fetches each page, saves the raw HTML, parses it, and appends rows to a CSV with a running position across pages:

def run(context, queries, out=OUT):
    page = context.new_page()
    new_file = not os.path.exists(out)
    with open(out, "a", newline="", encoding="utf-8") as fh:
        writer = csv.DictWriter(fh, fieldnames=FIELDS)
        if new_file:
            writer.writeheader()
        first = True
        for query in queries:
            position = 0
            for n in range(1, PAGES_PER_QUERY + 1):
                if not first:
                    time.sleep(random.uniform(*DELAY_SECONDS))
                first = False
                html = fetch_serp(page, search_url(query, n))
                ...

Choices worth copying:

  • Pace per IP, with jitter. 20 to 45 seconds between pages from one IP is a conservative start for a rank tracker. A fixed interval is itself a pattern; a random one is not.
  • Stop paginating early. A page with no results ends that query.
  • Keep raw HTML. One file per page in serp_raw/. It is your evidence when a rank looks wrong, and your test data when the parser needs fixing.
  • Store position, not just order. Rank is the reason you are here; store it with the query, country and timestamp so you can chart it.

For thousands of keywords, spread queries across several static IPs, one browser context per IP, rather than speeding up any single one.

The complete script to scrape Google search results

Save as serp.py:

import csv
import os
import random
import re
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlencode, urlsplit

from bs4 import BeautifulSoup
from playwright.sync_api import sync_playwright

QUERIES = ["web scraping proxies", "static isp proxy"]
HL, GL = "en", "us"
PAGES_PER_QUERY = 2
DELAY_SECONDS = (20, 45)
OUT = "serp_results.csv"
RAW_DIR = Path("serp_raw")
FIELDS = ["fetched_at", "query", "gl", "page", "position", "title", "url", "snippet"]


class Blocked(Exception):
    pass


def search_url(query, page):
    params = {"q": query, "hl": HL, "gl": GL}
    if page > 1:
        params["start"] = (page - 1) * 10
    return "https://www.google.com/search?" + urlencode(params)


def is_organic(href):
    host = urlsplit(href).hostname or ""
    return href.startswith("http") and not host.endswith(("google.com", "googleusercontent.com"))


def snippet_for(block, title):
    best = ""
    for el in block.find_all(["div", "span"]):
        if el.find("h3") or el.find("cite"):
            continue
        text = el.get_text(" ", strip=True)
        if title not in text and len(text) > len(best):
            best = text
    return best


def parse_results(html):
    soup = BeautifulSoup(html, "html.parser")
    results, seen = [], set()
    for h3 in soup.select("a[href] h3"):
        link = h3.find_parent("a")
        href = link["href"]
        if not is_organic(href) or href in seen:
            continue
        seen.add(href)
        block = link
        while block.parent is not None and len(block.parent.select("a[href] h3")) == 1:
            block = block.parent
        title = h3.get_text(" ", strip=True)
        results.append({"title": title, "url": href, "snippet": snippet_for(block, title)})
    return results


def check_page(url, html):
    if "/sorry/" in url or 'id="captcha-form"' in html:
        raise Blocked(f"CAPTCHA page at {url[:80]}")
    if "<h3" not in html:
        if "unusual traffic" in html:
            raise Blocked("unusual traffic page")
        if "/httpservice/retry/enablejs" in html:
            raise Blocked("Google served its enable-JavaScript page")


def accept_consent(page):
    button = page.get_by_role("button", name="Reject all", exact=True).first
    if urlsplit(page.url).hostname.startswith("consent.") or button.is_visible():
        button.click()
        page.wait_for_url(re.compile(r"^https://www\.google\.[a-z.]+/search"), timeout=15_000)
        page.wait_for_load_state("domcontentloaded")


def fetch_serp(page, url):
    page.goto(url, wait_until="domcontentloaded", timeout=45_000)
    accept_consent(page)
    try:
        page.wait_for_selector("a[href] h3", timeout=10_000)
    except Exception:
        pass
    html = page.content()
    check_page(page.url, html)
    return html


def run(context, queries, out=OUT):
    page = context.new_page()
    new_file = not os.path.exists(out)
    with open(out, "a", newline="", encoding="utf-8") as fh:
        writer = csv.DictWriter(fh, fieldnames=FIELDS)
        if new_file:
            writer.writeheader()
        first = True
        for query in queries:
            position = 0
            for n in range(1, PAGES_PER_QUERY + 1):
                if not first:
                    time.sleep(random.uniform(*DELAY_SECONDS))
                first = False
                html = fetch_serp(page, search_url(query, n))
                now = datetime.now(timezone.utc)
                RAW_DIR.mkdir(exist_ok=True)
                slug = re.sub(r"\W+", "-", query).strip("-")
                (RAW_DIR / f"{now:%Y%m%dT%H%M%S}-{slug}-p{n}.html").write_text(html, encoding="utf-8")
                results = parse_results(html)
                stamp = now.isoformat(timespec="seconds")
                for r in results:
                    position += 1
                    writer.writerow({"fetched_at": stamp, "query": query, "gl": GL,
                                     "page": n, "position": position, **r})
                print(f"{query!r} page {n}: {len(results)} results")
                if not results:
                    break


def main():
    proxy = {
        "server": f"http://{os.environ['PROXY_HOST']}:{os.environ['PROXY_PORT']}",
        "username": os.environ["PROXY_USER"],
        "password": os.environ["PROXY_PASS"],
    }
    with sync_playwright() as p:
        browser = p.chromium.launch(channel="chrome", proxy=proxy)
        context = browser.new_context(locale="en-US", viewport={"width": 1366, "height": 900})
        try:
            run(context, QUERIES)
        except Blocked as exc:
            sys.exit(f"stopped: {exc}. Pause this IP before trying again.")
        finally:
            browser.close()


if __name__ == "__main__":
    main()

How we ran it, on 2026-09-30 with Python 3.12, Playwright 1.63 and Chrome, through a local authenticating proxy: to stay within five live requests to Google, we drove run() with Playwright's request routing standing in for google.com. It served a consent dialog, then the saved results page twice, then the saved /sorry/ page. The run clicked through consent, wrote ten rows and two raw HTML files, and stopped with Blocked on the CAPTCHA page:

'web scraping proxies' page 1: 5 results
'web scraping proxies' page 2: 5 results
Blocked raised: CAPTCHA page at https://www.google.com/search?q=static+isp+proxy&hl=en&gl=us

The live requests, made with the same launch code plus the --disable-http2 flag described under troubleshooting, reached Google and got the /sorry/ page described in step 4.

Google search proxy: which IPs, and where

Google localises results by the gl parameter and by where the request comes from. For rank tracking, the IP's location is part of the measurement: a German keyword checked from a US address is a different ranking.

Proxy typeLocation choiceFit for Google SERP scraping
ISP (static)Choose by country, region, city and carrierLocal rankings from a consumer carrier's address; one steady identity per market
Datacenter (static)Choose by country, region and cityLow cost per IP; hosting ranges tend to meet challenges sooner, so pace them lower
Residential (rotating)None todaySpreads volume wide, but you cannot pick the country the ranking comes from

For a rank tracker, that makes static ISP IPs the first choice: pick the city, and every check for that market comes from the same place. They are sold one at a time from $3.20/IP a month, so a tracker covering three markets can start with three IPs. Our SERP tracking use case walks through a setup, and the US proxy locations page shows what "choose the city" looks like in practice.

Whatever the type, do the arithmetic from the target's side: pages per day divided by what one IP can send without meeting /sorry/ gives the number of IPs you need. Only you can measure that second number, for your keywords and your pace.

Troubleshooting

  • Navigation times out behind the proxy, but other sites load. Through our local test proxy, Chrome's HTTP/2 connection to Google stalled until we launched with args=["--disable-http2"]. If you see the same symptom, try that flag and tell your proxy provider.
  • Every page has zero results, no block. Save one page and open it. Usually it is a consent step in another language, or Google has changed where titles sit.
  • Snippets come back empty. Some result types (videos, sitelinks) have no text snippet. That is normal; keep the row.
  • Results do not match what you see at home. Signed-in history, hl, gl and the IP's location all change rankings. Compare like with like.

For the Python groundwork this page builds on (sessions, retries, parsing), see web scraping with proxies in Python.