A price monitoring scraper in Python does five jobs on a timer: fetch each product page (through proxies once you watch more than a handful), parse the price and stock status, normalise the price into a number and a currency, compare it with the last price you stored, and alert when it moves. This tutorial builds all five with requests, BeautifulSoup and SQLite, then schedules the job with cron or APScheduler.
We use the practice shop books.toscrape.com, which exists for scraping exercises, so every command here runs as written. Every Python script on this page ran on 2026-09-30 with Python 3.12, requests 2.34, beautifulsoup4 4.15 and APScheduler 3.11, through two local authenticating proxies.
How a price monitoring scraper in Python fits together
| Stage | What it does | Where it lives below |
|---|---|---|
| Config | Which products to watch and how to read each shop | PRODUCTS, SELECTORS |
| Fetch | Download each page through a proxy, politely | make_session, allowed_by_robots |
| Parse | Pull title, price text and stock text | scrape |
| Normalise | Turn "£51.77" into Decimal("51.77"), GBP | normalise_price |
| Store | Append every observation to SQLite | record |
| Compare and alert | Diff against the last observation, notify | changes, alert |
| Schedule | Run it every few hours | cron or APScheduler |
If you have not scraped with Python before, start with our web scraping with proxies in Python tutorial: it covers sessions, retries and the charset trap in more depth. This page assumes those basics and spends its time on what makes a monitor different, which is state.
Setup
python -m venv venv && . venv/bin/activate
pip install requests beautifulsoup4 apscheduler
The proxies come from an environment variable, one URL per static IP, comma-separated. On ProxyHive each ISP or datacenter IP is its own endpoint; open the order in the dashboard, copy each IP in USER:PASS@HOST:PORT form and put http:// in front:
export PROXY_LIST="http://USERNAME:PASSWORD@HOST1:PORT1,http://USERNAME:PASSWORD@HOST2:PORT2"
export WEBHOOK_URL="https://hooks.example.com/your-webhook"
WEBHOOK_URL is optional. Without it, alerts print to the console.
Step 1: describe the products you watch
Keep what you watch in data, not in code paths. A product is an ID you choose and a URL; a shop is a set of selectors keyed by hostname:
PRODUCTS = [
{"sku": "attic", "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"},
{"sku": "velvet", "url": "https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html"},
{"sku": "soumission", "url": "https://books.toscrape.com/catalogue/soumission_998/index.html"},
{"sku": "sharp-objects", "url": "https://books.toscrape.com/catalogue/sharp-objects_997/index.html"},
]
SELECTORS = {
"books.toscrape.com": {"title": "div.product_main h1", "price": "p.price_color", "stock": "p.availability"},
}
Adding a second shop means one new SELECTORS entry and its product URLs. When the list grows past a few dozen, move it to a CSV or a table in the same SQLite file.
Step 2: fetch product pages through proxies, politely
Each proxy gets its own requests.Session with a retry policy, and the run cycles through them so consecutive products leave from different IPs:
def make_session(proxy_url):
retry = Retry(total=3, backoff_factor=2, status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET"], respect_retry_after_header=True)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
session.headers.update({"User-Agent": USER_AGENT, "Accept-Language": "en-GB,en;q=0.9"})
session.proxies = {"http": proxy_url, "https": proxy_url}
return session
Politeness is three habits in this script:
- Read robots.txt once per shop.
urllib.robotparserdecides whether a URL may be fetched. A missing file (books.toscrape.com answers 404) means everything is allowed, so the code setsallow_allin that case. - Pause between products. A random 2 to 5 seconds on the practice site. On a real shop, lengthen it; a monitor has hours to finish, so there is no reason to hurry.
- Honour Retry-After. When a shop answers 429 with a wait time, urllib3 waits that long before retrying.
_robots = {}
def allowed_by_robots(session, url):
parts = urlsplit(url)
root = f"{parts.scheme}://{parts.netloc}"
if root not in _robots:
parser = RobotFileParser()
resp = session.get(f"{root}/robots.txt", timeout=(5, 15))
if resp.status_code >= 400:
parser.allow_all = True
else:
parser.parse(resp.text.splitlines())
_robots[root] = parser
return _robots[root].can_fetch(USER_AGENT, url)
For more than a couple of IPs, swap the simple cycle for a pool that benches failing proxies; rotating proxies in Python has one you can drop in.
Step 3: parse the price and availability
On the practice shop, the product page shows the price in p.price_color and the stock line in p.availability, for example "In stock (22 available)":
def scrape(session, product):
resp = session.get(product["url"], timeout=(5, 20))
resp.raise_for_status()
soup = BeautifulSoup(resp.content, "html.parser")
sel = SELECTORS[urlsplit(product["url"]).hostname]
price, currency = normalise_price(soup.select_one(sel["price"]).get_text())
stock_text = soup.select_one(sel["stock"]).get_text(" ", strip=True)
count = re.search(r"\((\d+) available\)", stock_text)
return Observation(
sku=product["sku"],
title=soup.select_one(sel["title"]).get_text(strip=True),
price=price,
currency=currency,
in_stock=stock_text.lower().startswith("in stock"),
stock_count=int(count.group(1)) if count else None,
)
If a selector stops matching, select_one returns None and .get_text() raises AttributeError. The run loop catches that and counts the product as failed rather than storing a blank price. That matters: a monitor that records "no price" as a price will alert you about a drop to nothing.
Real shops: read the structured data first
Many shops embed a schema.org Product with an Offer in a JSON-LD script tag, for search engines. When it is there, it is a steadier source than CSS classes, because the shop keeps it correct for Google. This helper returns the first product offer it finds, or None so you can fall back to selectors:
import json
from bs4 import BeautifulSoup
def offer_from_jsonld(html):
soup = BeautifulSoup(html, "html.parser")
for tag in soup.select('script[type="application/ld+json"]'):
try:
data = json.loads(tag.string or "")
except json.JSONDecodeError:
continue
items = data if isinstance(data, list) else data.get("@graph", [data])
for item in items:
if item.get("@type") != "Product":
continue
offer = item.get("offers", {})
offer = offer[0] if isinstance(offer, list) else offer
return {
"price": offer.get("price") or offer.get("lowPrice"),
"currency": offer.get("priceCurrency"),
"in_stock": str(offer.get("availability", "")).endswith("InStock"),
}
return None
The practice shop has no JSON-LD, so we tested this on sample documents: a @graph with a product at 39.99 EUR in stock, a list with a product out of stock, and a page with none. All three came back as expected. Check the schema.org Offer definition for the other fields shops fill in, such as priceValidUntil.
Step 4: normalise prices and currencies
Prices arrive as text in local formats: "£51.77", "€1.299,00", "$1,299.99", "1 299,50 €". Store them as Decimal, never float, because floats cannot represent most prices exactly and a comparison between two of them can report a change that never happened.
CURRENCIES = {"£": "GBP", "€": "EUR", "$": "USD", "¥": "JPY"}
def normalise_price(text):
text = text.strip()
currency = next((code for sym, code in CURRENCIES.items() if sym in text), None)
if currency is None:
found = re.search(r"\b[A-Z]{3}\b", text)
currency = found.group(0) if found else "UNKNOWN"
number = re.sub(r"[^\d.,]", "", text)
if "," in number and "." in number:
decimal_mark = "," if number.rfind(",") > number.rfind(".") else "."
elif "," in number:
decimal_mark = "," if len(number.rsplit(",", 1)[1]) != 3 else ""
else:
decimal_mark = "."
if decimal_mark == ",":
number = number.replace(".", "").replace(",", ".")
else:
number = number.replace(",", "")
return Decimal(number), currency
The rule: when both separators appear, the last one is the decimal mark; a lone comma followed by exactly three digits is a thousands separator. Our test inputs came back like this:
'£51.77' (Decimal('51.77'), 'GBP')
'€1.299,00' (Decimal('1299.00'), 'EUR')
'$1,299.99' (Decimal('1299.99'), 'USD')
'1 299,50 €' (Decimal('1299.50'), 'EUR')
'USD 12.50' (Decimal('12.50'), 'USD')
'¥1,200' (Decimal('1200'), 'JPY')
'12,5 EUR' (Decimal('12.5'), 'EUR')
One case no rule can settle: "€1.299" could be one euro and 29.9 cents or one thousand two hundred and ninety-nine euros. If a shop formats prices that way, set its decimal mark in SELECTORS rather than guessing. And "$" is ambiguous between US, Canadian and Australian dollars; map it per shop for anything beyond a demo. Do not convert currencies at scrape time either. Store what the shop showed, and convert in reporting with a dated rate.
Step 5: store the price history in SQLite
A monitor is only as useful as its history. One table, one row per observation, with the price stored as text so the Decimal survives the round trip exactly:
def open_db(path=DB_PATH):
db = sqlite3.connect(path)
db.execute("""
CREATE TABLE IF NOT EXISTS observations (
id INTEGER PRIMARY KEY,
sku TEXT NOT NULL,
observed_at TEXT NOT NULL,
title TEXT NOT NULL,
price TEXT NOT NULL,
currency TEXT NOT NULL,
in_stock INTEGER NOT NULL,
stock_count INTEGER
)""")
db.execute("CREATE INDEX IF NOT EXISTS obs_sku_time ON observations (sku, observed_at)")
return db
Every run appends, even when nothing changed. That costs a few bytes per product per run and gives you something a changes-only log cannot: proof that you checked. When a competitor says their price "never" dropped, a gap in your data and an unchanged row are different answers. Timestamps are ISO 8601 in UTC, so they sort as text. SQLite ships with Python as the sqlite3 module, so there is no server to run.
Step 6: detect changes and send alerts
Compare each new observation with the last stored one for the same product, before recording the new one:
def changes(previous, obs):
if previous is None:
return []
old_price, old_currency, old_stock = Decimal(previous[0]), previous[1], bool(previous[2])
found = []
if old_currency == obs.currency and old_price != obs.price:
pct = (obs.price - old_price) / old_price * 100
found.append(f"price {old_price} -> {obs.price} {obs.currency} ({pct:+.1f}%)")
if old_stock != obs.in_stock:
found.append("back in stock" if obs.in_stock else "out of stock")
return found
def alert(obs, found):
message = f"[{obs.sku}] {obs.title}: " + "; ".join(found)
print("ALERT", message)
webhook = os.environ.get("WEBHOOK_URL")
if webhook:
requests.post(webhook, data=json.dumps({"text": message}),
headers={"Content-Type": "application/json"}, timeout=10)
The webhook payload is {"text": ...}, which Slack incoming webhooks accept as is; for Discord, rename the key to content. The alert goes out directly, not through the proxy: it is your own endpoint.
A first observation never alerts; there is nothing to compare it with. A currency change does not alert as a price change either, because comparing pounds with euros is meaningless. It is also a sign the shop is now showing you another market, which usually means the IP's location changed.
The complete price monitoring script
Save as price_monitor.py:
import json
import os
import random
import re
import sqlite3
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from decimal import Decimal
from itertools import cycle
from urllib.parse import urlsplit
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
PRODUCTS = [
{"sku": "attic", "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"},
{"sku": "velvet", "url": "https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html"},
{"sku": "soumission", "url": "https://books.toscrape.com/catalogue/soumission_998/index.html"},
{"sku": "sharp-objects", "url": "https://books.toscrape.com/catalogue/sharp-objects_997/index.html"},
]
SELECTORS = {
"books.toscrape.com": {"title": "div.product_main h1", "price": "p.price_color", "stock": "p.availability"},
}
DB_PATH = "prices.db"
DELAY_SECONDS = (2, 5)
USER_AGENT = "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36"
CURRENCIES = {"£": "GBP", "€": "EUR", "$": "USD", "¥": "JPY"}
@dataclass
class Observation:
sku: str
title: str
price: Decimal
currency: str
in_stock: bool
stock_count: int | None
def normalise_price(text):
text = text.strip()
currency = next((code for sym, code in CURRENCIES.items() if sym in text), None)
if currency is None:
found = re.search(r"\b[A-Z]{3}\b", text)
currency = found.group(0) if found else "UNKNOWN"
number = re.sub(r"[^\d.,]", "", text)
if "," in number and "." in number:
decimal_mark = "," if number.rfind(",") > number.rfind(".") else "."
elif "," in number:
decimal_mark = "," if len(number.rsplit(",", 1)[1]) != 3 else ""
else:
decimal_mark = "."
if decimal_mark == ",":
number = number.replace(".", "").replace(",", ".")
else:
number = number.replace(",", "")
return Decimal(number), currency
def proxy_urls():
listed = os.environ.get("PROXY_LIST", "")
return [u.strip() for u in listed.split(",") if u.strip()]
def make_session(proxy_url):
retry = Retry(total=3, backoff_factor=2, status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET"], respect_retry_after_header=True)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
session.headers.update({"User-Agent": USER_AGENT, "Accept-Language": "en-GB,en;q=0.9"})
session.proxies = {"http": proxy_url, "https": proxy_url}
return session
_robots = {}
def allowed_by_robots(session, url):
parts = urlsplit(url)
root = f"{parts.scheme}://{parts.netloc}"
if root not in _robots:
parser = RobotFileParser()
resp = session.get(f"{root}/robots.txt", timeout=(5, 15))
if resp.status_code >= 400:
parser.allow_all = True
else:
parser.parse(resp.text.splitlines())
_robots[root] = parser
return _robots[root].can_fetch(USER_AGENT, url)
def scrape(session, product):
resp = session.get(product["url"], timeout=(5, 20))
resp.raise_for_status()
soup = BeautifulSoup(resp.content, "html.parser")
sel = SELECTORS[urlsplit(product["url"]).hostname]
price, currency = normalise_price(soup.select_one(sel["price"]).get_text())
stock_text = soup.select_one(sel["stock"]).get_text(" ", strip=True)
count = re.search(r"\((\d+) available\)", stock_text)
return Observation(
sku=product["sku"],
title=soup.select_one(sel["title"]).get_text(strip=True),
price=price,
currency=currency,
in_stock=stock_text.lower().startswith("in stock"),
stock_count=int(count.group(1)) if count else None,
)
def open_db(path=DB_PATH):
db = sqlite3.connect(path)
db.execute("""
CREATE TABLE IF NOT EXISTS observations (
id INTEGER PRIMARY KEY,
sku TEXT NOT NULL,
observed_at TEXT NOT NULL,
title TEXT NOT NULL,
price TEXT NOT NULL,
currency TEXT NOT NULL,
in_stock INTEGER NOT NULL,
stock_count INTEGER
)""")
db.execute("CREATE INDEX IF NOT EXISTS obs_sku_time ON observations (sku, observed_at)")
return db
def last_observation(db, sku):
return db.execute(
"SELECT price, currency, in_stock FROM observations WHERE sku = ? ORDER BY observed_at DESC, id DESC LIMIT 1",
(sku,),
).fetchone()
def record(db, obs):
db.execute(
"INSERT INTO observations (sku, observed_at, title, price, currency, in_stock, stock_count) VALUES (?, ?, ?, ?, ?, ?, ?)",
(obs.sku, datetime.now(timezone.utc).isoformat(timespec="seconds"), obs.title,
str(obs.price), obs.currency, int(obs.in_stock), obs.stock_count),
)
db.commit()
def changes(previous, obs):
if previous is None:
return []
old_price, old_currency, old_stock = Decimal(previous[0]), previous[1], bool(previous[2])
found = []
if old_currency == obs.currency and old_price != obs.price:
pct = (obs.price - old_price) / old_price * 100
found.append(f"price {old_price} -> {obs.price} {obs.currency} ({pct:+.1f}%)")
if old_stock != obs.in_stock:
found.append("back in stock" if obs.in_stock else "out of stock")
return found
def alert(obs, found):
message = f"[{obs.sku}] {obs.title}: " + "; ".join(found)
print("ALERT", message)
webhook = os.environ.get("WEBHOOK_URL")
if webhook:
requests.post(webhook, data=json.dumps({"text": message}),
headers={"Content-Type": "application/json"}, timeout=10)
def run_once():
proxies = proxy_urls()
if not proxies:
raise SystemExit("set PROXY_LIST to one or more http://USERNAME:PASSWORD@HOST:PORT URLs")
sessions = cycle([make_session(p) for p in proxies])
db = open_db()
ok = failed = 0
for i, product in enumerate(PRODUCTS):
if i:
time.sleep(random.uniform(*DELAY_SECONDS))
session = next(sessions)
try:
if not allowed_by_robots(session, product["url"]):
print(f"skip {product['sku']}: disallowed by robots.txt")
continue
obs = scrape(session, product)
except (requests.RequestException, AttributeError) as exc:
failed += 1
print(f"fail {product['sku']}: {type(exc).__name__}: {exc}")
continue
found = changes(last_observation(db, obs.sku), obs)
record(db, obs)
ok += 1
if found:
alert(obs, found)
db.close()
print(f"{datetime.now():%H:%M:%S} checked {ok} products, {failed} failed")
if __name__ == "__main__":
run_once()
What happened when we ran it
The practice shop's prices never change, so to see an alert we ran the script, then edited the stored rows to pretend "A Light in the Attic" had cost £55.00 and "Tipping the Velvet" had been out of stock, and ran it again. A local HTTP server stood in for the webhook:
$ python price_monitor.py
01:30:22 checked 4 products, 0 failed
$ python price_monitor.py
ALERT [attic] A Light in the Attic: price 55.00 -> 51.77 GBP (-5.9%)
ALERT [velvet] Tipping the Velvet: back in stock
01:30:36 checked 4 products, 0 failed
The webhook received both messages as JSON. With a wrong proxy password, all four products failed with ProxyError (the proxy answered 407), the script reported checked 0 products, 4 failed, and it stored nothing, which is the behaviour you want: a failed check leaves a gap, never a false price.
To read the history back:
import sqlite3
db = sqlite3.connect("prices.db")
rows = db.execute("""
SELECT sku, observed_at, price, currency, in_stock
FROM observations
WHERE sku = ?
ORDER BY observed_at
""", ("attic",))
for row in rows:
print(row)
Schedule the monitor: cron or APScheduler
cron is the simplest if the machine is always on. Four runs a day, at minute 17 so you are not in the crowd that runs everything on the hour:
17 */6 * * * cd /opt/prices && PROXY_LIST="http://USERNAME:PASSWORD@HOST:PORT" ./venv/bin/python price_monitor.py >> monitor.log 2>&1
APScheduler keeps the schedule in Python, which suits a container or a process manager. This is the 3.x API (APScheduler 4 changes it):
import os
from apscheduler.schedulers.blocking import BlockingScheduler
from price_monitor import run_once
scheduler = BlockingScheduler(timezone="UTC")
scheduler.add_job(
run_once,
"interval",
minutes=int(os.environ.get("CHECK_EVERY_MINUTES", "360")),
jitter=60,
max_instances=1,
coalesce=True,
)
if __name__ == "__main__":
run_once()
scheduler.start()
jitter moves each run by up to a minute so the monitor has no exact rhythm, max_instances=1 stops a slow run from overlapping the next, and coalesce=True runs a missed job once rather than several times in a row after a pause. We ran it with CHECK_EVERY_MINUTES=1 for three minutes and got a run at start and another a minute and a half later. The APScheduler user guide covers cron-style triggers and persistent job stores.
Which proxies for price monitoring
Many shops show prices, stock and even catalogues by the visitor's country. A monitor that reads a German shop from a US address is measuring something no German customer sees. So location is the first requirement, and stability is the second: the same market should be read from the same place every run.
| Need | Proxy type | Why |
|---|---|---|
| Local prices in a few markets | ISP (static), chosen by country, region, city and carrier | Consumer-carrier addresses, one steady IP per market |
| Many products on lenient shops | Datacenter (static), chosen by country, region and city | Low cost per IP, fine where the shop does not check network type |
| Wide sweeps where location does not matter | Residential (rotating), per GB | Spreads load widely; no location choice today |
Static IPs are sold one at a time, from $3.20/IP a month for a single ISP IP, so a monitor for three markets can start with three. Our price monitoring use case covers the approach in more detail, and the Germany locations page shows the marketplaces people watch there. Price monitoring of public product pages is an allowed use on our network; the allowed-use policy lists the few retail cases, such as limited-release drops, where we ask you to talk to us first.
Taking it further
- JavaScript-rendered prices. If the price is not in the HTML, look in the browser's network tab for the JSON the page loads, and call that. Only reach for Playwright when there is no such endpoint.
- Many products. Past a few thousand pages per run, move fetching to a thread pool or Scrapy and keep this script's storage and comparison logic.
- Search visibility. Pair price with where the product ranks: our guide to scraping Google search results builds the rank-tracking half.