Web scraping with Node.js takes four packages: undici to fetch pages (its ProxyAgent sends each request through a proxy), cheerio to parse the HTML with CSS selectors, p-limit to cap how many requests run at once, and Playwright for the pages that only exist after JavaScript runs. Add retries with exponential backoff and a rotation over your proxy IPs, and you have a scraper that finishes the job and tells you what it missed.

This tutorial builds that scraper step by step and crawls all 1,000 books from the practice site books.toscrape.com into CSV and JSON, then scrapes a JavaScript-rendered practice page with Playwright. Every script here ran on 2026-09-30 on Node 22.23 with undici 8.11, cheerio 1.2, p-limit 7.3 and Playwright 1.63, through two local authenticating proxies.

What you need

  1. Node.js 22 (p-limit 7 and undici 8 need a current Node; ES modules throughout).
  2. A project with the packages:
mkdir books-scraper && cd books-scraper
npm init -y && npm pkg set type=module
npm install undici cheerio p-limit playwright
  1. Proxies. Each ISP or datacenter IP on a ProxyHive order is its own endpoint. Open the order in the dashboard, copy each IP in USER:PASS@HOST:PORT form, add http:// and join them with commas:
export PROXY_LIST="http://USERNAME:PASSWORD@HOST1:PORT1,http://USERNAME:PASSWORD@HOST2:PORT2"

Use the HTTP port from the order. Our Node.js proxy setup guide covers axios, SOCKS5 and environment-variable proxies; this page stays on the scraping itself.

Step 1: fetch a page through a proxy with undici

import { fetch, ProxyAgent } from 'undici';

const dispatcher = new ProxyAgent(process.env.PROXY_LIST.split(',')[0]);
const res = await fetch('https://books.toscrape.com/', {
  dispatcher,
  signal: AbortSignal.timeout(20_000),
});
console.log(res.status, (await res.text()).length);

Import fetch from undici, not the global one. Node 22 bundles an older undici internally, and handing an npm ProxyAgent to the global fetch fails with invalid onRequestStart method. AbortSignal.timeout matters too: fetch has no overall timeout of its own, and one stalled connection can hold a slot in your crawl indefinitely.

A small win over Python's requests: res.text() always decodes as UTF-8, as the Fetch standard says, so the pound signs on this site arrive intact.

Step 2: parse the HTML with cheerio

cheerio gives you jQuery-style selectors over an HTML string, with no browser. Each book on a listing page is an article.product_pod:

function parseListing(html, pageUrl) {
  const $ = cheerio.load(html);
  const books = $('article.product_pod')
    .map((_, el) => {
      const link = $(el).find('h3 a');
      return {
        title: link.attr('title'),
        price: Number($(el).find('p.price_color').text().replace(/[^\d.]/g, '')),
        currency: 'GBP',
        rating: RATINGS[$(el).find('p.star-rating').attr('class').split(' ')[1]] ?? null,
        inStock: $(el).find('p.availability').text().includes('In stock'),
        url: new URL(link.attr('href'), pageUrl).href,
      };
    })
    .get();
  const total = Number($('li.current').text().match(/of (\d+)/)?.[1] ?? 1);
  return { books, total };
}

Three habits in there worth keeping:

  • Resolve links with new URL(href, pageUrl). Listing pages use relative links; storing them raw gives you URLs that only work on the page they came from.
  • Turn prices into numbers at parse time, and record the currency next to them.
  • Read the page count from the page ("Page 1 of 50") instead of hard-coding it, so the crawl adapts when the catalogue grows.

Step 3: retry with exponential backoff and Retry-After

Networks fail and servers push back. The retry loop below treats three cases differently:

  1. Network and proxy errors (undici throws TypeError: fetch failed): retry.
  2. 429 and 5xx responses: retry, waiting as long as Retry-After asks when the server sends it.
  3. Other statuses, such as 404: fail at once. Retrying a missing page wastes a request.
async function getHtml(url, { attempts = 4 } = {}) {
  const start = turn++;
  for (let attempt = 1; ; attempt++) {
    const dispatcher = agents[(start + attempt - 1) % agents.length];
    try {
      const res = await fetch(url, { dispatcher, headers: HEADERS, signal: AbortSignal.timeout(20_000) });
      if (res.ok) return await res.text();
      await res.body?.cancel();
      throw new HttpError(res.status, url, Number(res.headers.get('retry-after')) || 0);
    } catch (err) {
      const retryable = !(err instanceof HttpError) || RETRY_STATUS.has(err.status);
      if (!retryable || attempt >= attempts) throw err;
      const backoff = err.retryAfter ? err.retryAfter * 1000 : 2 ** attempt * 500 + Math.random() * 500;
      console.warn(`retry ${attempt} for ${url} in ${Math.round(backoff)} ms (${err.cause?.message ?? err.message})`);
      await sleep(backoff);
    }
  }
}

The backoff doubles each time (about 1, 2 and 4 seconds) with up to half a second of random jitter, so parallel requests that failed together do not retry together. res.body?.cancel() releases the connection for responses we do not read.

We tested it against a local server that answered 429 with Retry-After: 1, then 503, then 200: the function waited 1,000 ms, then about 2,100 ms, and returned the page on the third request. Against a 404 it gave up after one request.

When a request fails, log err.cause. undici's own message is always fetch failed; the cause holds the reason. With a wrong proxy password ours read Proxy response (407) !== 200 when HTTP Tunneling. Our proxy error codes guide explains the rest.

Step 4: rotate requests across a list of static IPs

Create one ProxyAgent per IP, once, and reuse them: each agent keeps its own pool of open connections to its proxy.

const proxyUrls = (process.env.PROXY_LIST ?? '').split(',').map((s) => s.trim()).filter(Boolean);
const agents = proxyUrls.map((url) => new ProxyAgent(url));
let turn = 0;

getHtml takes a starting position from a shared counter (turn++), which spreads requests round-robin over the IPs, and each retry moves one step along from there. That second detail matters. Our first version drew the next agent from the shared counter on every attempt, and with a broken proxy in a list of two, several pages drew the broken one twice in a row. Walking forward from the request's own starting point guarantees that a retry uses a different IP. With one proxy deliberately given the wrong password, the finished crawl still saved all 1,000 books; the retries showed up as warnings, and no page failed.

This is the simplest useful rotation. The ideas in rotating proxies in Python, such as benching a failing IP for a cooldown and keeping one IP per domain, port straight to this structure.

Step 5: cap concurrency with p-limit

Promise.all over 49 fetches starts all 49 at once. p-limit queues them behind a limit:

const limit = pLimit(PER_PROXY * agents.length);
const pages = await Promise.all(
  Array.from({ length: total - 1 }, (_, i) => i + 2).map((n) =>
    limit(async () => {
      const url = `${BASE}page-${n}.html`;
      try {
        return parseListing(await getHtml(url), url).books;
      } catch (err) {
        failed.push(n);
        console.error(`page ${n} failed: ${err.message}`);
        return [];
      }
    }),
  ),
);

The limit is PER_PROXY times the number of IPs, because concurrency per IP is the number a target's rate limiter sees. Two per IP is gentle; raise it while responses stay fast and clean, and lower it at the first 429. Each page catches its own error, so one bad page is recorded in failed and the rest of the crawl continues.

Step 6: save the results to CSV and JSON

JSON is one line. CSV needs quoting for titles with commas or quotes, which this catalogue has plenty of:

const toCsv = (rows) => {
  const cols = Object.keys(rows[0]);
  const cell = (v) => (/[",\n]/.test(String(v)) ? `"${String(v).replace(/"/g, '""')}"` : String(v));
  return [cols.join(','), ...rows.map((r) => cols.map((c) => cell(r[c])).join(','))].join('\n') + '\n';
};

Sort the rows before writing. Pages finish in whatever order the network allows, and sorted output lets you diff two runs.

The complete Node.js scraper

Save as crawl.mjs:

import { writeFile } from 'node:fs/promises';
import { setTimeout as sleep } from 'node:timers/promises';

import * as cheerio from 'cheerio';
import pLimit from 'p-limit';
import { fetch, ProxyAgent } from 'undici';

const BASE = 'https://books.toscrape.com/catalogue/';
const PER_PROXY = 2;
const HEADERS = {
  'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36',
  'accept-language': 'en-GB,en;q=0.9',
};
const RATINGS = { One: 1, Two: 2, Three: 3, Four: 4, Five: 5 };
const RETRY_STATUS = new Set([429, 500, 502, 503, 504]);

const proxyUrls = (process.env.PROXY_LIST ?? '').split(',').map((s) => s.trim()).filter(Boolean);
if (proxyUrls.length === 0) throw new Error('Set PROXY_LIST to http://USERNAME:PASSWORD@HOST:PORT[,...]');
const agents = proxyUrls.map((url) => new ProxyAgent(url));
let turn = 0;

class HttpError extends Error {
  constructor(status, url, retryAfter) {
    super(`${status} from ${url}`);
    this.status = status;
    this.retryAfter = retryAfter;
  }
}

async function getHtml(url, { attempts = 4 } = {}) {
  const start = turn++;
  for (let attempt = 1; ; attempt++) {
    const dispatcher = agents[(start + attempt - 1) % agents.length];
    try {
      const res = await fetch(url, { dispatcher, headers: HEADERS, signal: AbortSignal.timeout(20_000) });
      if (res.ok) return await res.text();
      await res.body?.cancel();
      throw new HttpError(res.status, url, Number(res.headers.get('retry-after')) || 0);
    } catch (err) {
      const retryable = !(err instanceof HttpError) || RETRY_STATUS.has(err.status);
      if (!retryable || attempt >= attempts) throw err;
      const backoff = err.retryAfter ? err.retryAfter * 1000 : 2 ** attempt * 500 + Math.random() * 500;
      console.warn(`retry ${attempt} for ${url} in ${Math.round(backoff)} ms (${err.cause?.message ?? err.message})`);
      await sleep(backoff);
    }
  }
}

function parseListing(html, pageUrl) {
  const $ = cheerio.load(html);
  const books = $('article.product_pod')
    .map((_, el) => {
      const link = $(el).find('h3 a');
      return {
        title: link.attr('title'),
        price: Number($(el).find('p.price_color').text().replace(/[^\d.]/g, '')),
        currency: 'GBP',
        rating: RATINGS[$(el).find('p.star-rating').attr('class').split(' ')[1]] ?? null,
        inStock: $(el).find('p.availability').text().includes('In stock'),
        url: new URL(link.attr('href'), pageUrl).href,
      };
    })
    .get();
  const total = Number($('li.current').text().match(/of (\d+)/)?.[1] ?? 1);
  return { books, total };
}

const toCsv = (rows) => {
  const cols = Object.keys(rows[0]);
  const cell = (v) => (/[",\n]/.test(String(v)) ? `"${String(v).replace(/"/g, '""')}"` : String(v));
  return [cols.join(','), ...rows.map((r) => cols.map((c) => cell(r[c])).join(','))].join('\n') + '\n';
};

const first = `${BASE}page-1.html`;
const { books, total } = parseListing(await getHtml(first), first);
console.log(`page 1: ${books.length} books, ${total} pages in total`);

const limit = pLimit(PER_PROXY * agents.length);
const failed = [];
const pages = await Promise.all(
  Array.from({ length: total - 1 }, (_, i) => i + 2).map((n) =>
    limit(async () => {
      const url = `${BASE}page-${n}.html`;
      try {
        return parseListing(await getHtml(url), url).books;
      } catch (err) {
        failed.push(n);
        console.error(`page ${n} failed: ${err.message}`);
        return [];
      }
    }),
  ),
);

const all = [...books, ...pages.flat()].sort((a, b) => a.url.localeCompare(b.url));
await writeFile('books.json', JSON.stringify(all, null, 2));
await writeFile('books.csv', toCsv(all));
console.log(`saved ${all.length} books; failed pages: ${JSON.stringify(failed)}`);
await Promise.all(agents.map((a) => a.close()));

Run it:

node crawl.mjs
page 1: 20 books, 50 pages in total
saved 1000 books; failed pages: []

books.json holds objects like this:

{
  "title": "10-Day Green Smoothie Cleanse: Lose Up to 15 Pounds in 10 Days!",
  "price": 49.71,
  "currency": "GBP",
  "rating": 5,
  "inStock": true,
  "url": "https://books.toscrape.com/catalogue/10-day-green-smoothie-cleanse-lose-up-to-15-pounds-in-10-days_581/index.html"
}

With every proxy given a wrong password, the first page failed after four attempts and the script exited with the 407 in the error's cause. Failing loudly before the crawl starts is the right outcome: nothing downstream should run on an empty file.

Going deeper: enrich each book from its product page

Listing pages carry the basics. The product page adds the UPC, the exact stock count and the category, which is what a catalogue or price database needs as a key. The crawl already has every product URL, so enrichment is one more pass through the same getHtml and the same limit. Add this to the end of crawl.mjs, before the agents are closed:

function parseDetail(html) {
  const $ = cheerio.load(html);
  const table = Object.fromEntries(
    $('table.table-striped tr').map((_, tr) => [[$(tr).find('th').text(), $(tr).find('td').text()]]).get(),
  );
  return {
    upc: table.UPC,
    available: Number(table.Availability?.match(/\d+/)?.[0] ?? 0),
    category: $('ul.breadcrumb li').eq(2).text().trim(),
  };
}

const detailed = await Promise.all(
  all.slice(0, 20).map((book) =>
    limit(async () => ({ ...book, ...parseDetail(await getHtml(book.url)) })),
  ),
);
console.log(detailed.slice(0, 2));

We ran it for the first 20 books; the first came back with upc: '96aa539bfd4c07e2', available: 10 and category: 'Health'. Two details make it hold up. The product table is read into an object keyed by its header cells, so a new row or a reordered table does not shift the values. And .map() returns a nested array on purpose: cheerio flattens one level of what the callback returns, so each pair has to be wrapped once more to survive as a pair.

Drop the .slice(0, 20) for the full catalogue. That is 1,000 more requests, so this is where the per-IP limit and the retry loop earn their keep. Keep the same PER_PROXY value, and let the number of IPs, not the per-IP rate, decide how fast the whole job goes.

Node.js or Python for scraping?

Both do the job, and the choice usually follows the rest of your stack. Node's strengths are async I/O by default (thousands of pending requests cost little), one language across scraper and browser code, and first-class Playwright and Puppeteer support. Python's are the data side: pandas, a mature Scrapy framework and more scraping examples in the wild. If your team already runs Node services, stay in Node; the patterns on this page mirror our Python scraping tutorial step for step, so moving between the two is mostly syntax.

Scraping JavaScript pages with Playwright in Node.js

cheerio reads what the server sends and never runs a script. The practice page quotes.toscrape.com/js builds its quotes in the browser: the HTML that fetch receives contains no div.quote elements at all (we checked). For pages like that, drive a browser. Playwright takes the proxy at launch, with credentials as separate fields:

import { writeFile } from 'node:fs/promises';

import { chromium } from 'playwright';

const proxy = new URL(process.env.PROXY_URL);
const browser = await chromium.launch({
  channel: 'chrome',
  proxy: {
    server: `http://${proxy.host}`,
    username: decodeURIComponent(proxy.username),
    password: decodeURIComponent(proxy.password),
  },
});
const page = await browser.newPage();
const quotes = [];

try {
  let url = 'https://quotes.toscrape.com/js/';
  while (url) {
    await page.goto(url, { waitUntil: 'domcontentloaded' });
    await page.locator('div.quote').first().waitFor();
    quotes.push(
      ...(await page.locator('div.quote').evaluateAll((els) =>
        els.map((el) => ({
          text: el.querySelector('span.text').textContent,
          author: el.querySelector('small.author').textContent,
          tags: [...el.querySelectorAll('a.tag')].map((a) => a.textContent),
        })),
      )),
    );
    const next = page.locator('li.next a');
    url = (await next.count()) ? new URL(await next.getAttribute('href'), page.url()).href : null;
  }
} finally {
  await browser.close();
}

await writeFile('quotes.json', JSON.stringify(quotes, null, 2));
console.log(`saved ${quotes.length} quotes, first by ${quotes[0].author}`);
saved 100 quotes, first by Albert Einstein

Notes from the run:

  • Wait for the data, not the page. waitFor() on the first div.quote returns once the script has rendered it. Waiting for load alone can leave you parsing an empty container.
  • Follow the "Next" link rather than guessing page numbers; the loop ends when there is none.
  • Chrome or bundled Chromium. We used installed Chrome (channel: 'chrome'). Run npx playwright install chromium and remove the channel line to use Playwright's own build.
  • HTTP port only. Chromium cannot send a username and password to a SOCKS5 proxy, so use the order's HTTP port.

A browser costs far more CPU, memory and traffic per page than fetch. Before reaching for one, open the browser's network tab: JavaScript pages usually load their data from a JSON endpoint that getHtml can call directly. Our Playwright proxy guide covers per-context proxies for several IPs in one browser.

Which proxies for web scraping with Node.js

The code does not change with the proxy type; the target decides.

Proxy typeUse it whenBilling
Datacenter (static)The target does not check network type: practice sites, APIs, many cataloguesPer IP, from $3.20/IP a month for one
ISP (static)The target refuses hosting ranges but tolerates a steady identityPer IP
Residential (rotating)Strict targets and wide sweeps; one endpoint, the provider rotates the exitPer GB; no location choice today

Start with a couple of datacenter proxies, measure clean pages against blocks over a few hundred requests, and move up only when the numbers say so. When the crawler outgrows a single script (request queues, persistence, autoscaling), Crawlee builds on these same pieces.

Scrape public data at a pace the site can absorb, and read the target's terms. What our network may be used for is public in the allowed-use policy. This is not legal advice.

Troubleshooting

SymptomCauseFix
invalid onRequestStart methodGlobal fetch with an npm undici ProxyAgentImport fetch from undici
TypeError: fetch failedNetwork or proxy error, reason hiddenLog err.cause; a 407 means wrong proxy credentials
pLimit is not a functionp-limit is ES modules only; on Node 22, require() returns its namespace objectUse import, or require('p-limit').default
Crawl hangs near the endA request with no timeoutPass AbortSignal.timeout() on every fetch
Empty results from cheerioThe page renders with JavaScriptFind the JSON endpoint, or use Playwright