Multi-Channel Selling

Ecommerce Competitor Price Monitoring: How to Build a Custom Scraper That Lasts

Ecommerce competitor price monitoring with a custom scraper: which prices to capture, how to extract and match them, store history, and respect robots.txt.

Long Nguyen Avatar

Long Nguyen

Fullstack Developer · AI Engineer · Researcher

• • 4 min read •

What a Competitor Price Monitoring Scraper Has to Do

A custom scraper for ecommerce competitor price monitoring is a pipeline, not a script. It fetches competitor product pages, extracts price and availability, matches each page to one of your SKUs, stores every observation, and tells you when something changed or broke.

Fetching and parsing is the easy part. Whether the data can be trusted depends on matching, history and failure detection:

  1. Fetch product URLs politely, on a schedule.
  2. Extract price, currency, list price, availability and shipping.
  3. Match each competitor URL to your own SKU.
  4. Store every observation, not just the latest price.
  5. Alert on meaningful price moves and on scraper failures.

The rest of this guide follows that order.

Custom Scraper vs Price Monitoring Tool: When Building Wins

A hosted monitoring tool is usually the faster start. A custom scraper wins when your problem does not fit a tool's assumptions.

Factor Custom scraper Hosted monitoring tool
Competitor sites Many niche, regional or oddly built sites A few mainstream retailers the tool already covers
Product matching Your rules: GTIN, MPN, internal part numbers, bundles The tool's matching, with limited control
Integration Straight into your database, catalog or repricing rules Export or API within the tool's limits
History and ownership Full history in your own database Depends on plan and retention
Cost shape Engineering time up front, then maintenance Recurring subscription, little engineering
Maintenance Yours: layout changes, blocks, monitoring The vendor's
Time to first data Longer Shorter

A sensible split for many sellers: use a tool for the handful of big competitors it handles well, and a scraper for the long tail of niche sites and for anything that must feed your own pricing logic.

If you decide to build and would rather not run the crawler yourself, Netalith's custom software service covers data crawling and processing, including price collection pipelines like the one described here.

Which Prices and Fields to Capture

Do not store one number called "price". Google's merchant listing documentation distinguishes three kinds of price: the active price, a strikethrough price shown during a sale, and a member price for a loyalty program. It encodes them in Offer and UnitPriceSpecification markup, using priceType for the strikethrough price and validForMemberTier for the member price. Mixing them is a reliable way to produce a wrong comparison.

Field Why it matters Typical source
Active price What a shopper pays right now price, or a price specification without priceType
Strikethrough price Separates a real cut from a permanent "was" price Price specification with priceType set to StrikethroughPrice
Member price A loyalty price is not the public price validForMemberTier
Currency Never compare raw numbers across currencies priceCurrency
Availability An out-of-stock rival is not a real price threat availability
Shipping cost and delivery time Needed to compare landed price shippingDetails, or visible page text
Sale window Tells you when a promotion ends validFrom, priceValidUntil, validThrough
Unit price Compares per-100 ml or per-kilo, not per-pack referenceQuantity

Two practical notes. When a store sells in several currencies, Google's guidance is one URL per currency, so key your observations on URL plus currency. And markup can lag the visible page or leave out tax, so audit a sample of extracted prices against what a shopper actually sees.

How to Extract Prices: Structured Data First, HTML Second

Try extraction methods from cheapest and most stable to most expensive:

  1. Product structured data (JSON-LD or microdata) in the page source. Stable, cheap, and often carries GTIN and MPN too.
  2. Site-specific CSS or XPath selectors for sites without usable markup.
  3. A rendered page in a headless browser, only when the price is injected by JavaScript.

Google recommends putting Product markup in the initial HTML, and warns that JavaScript-generated markup can make Shopping crawls less frequent and less reliable for fast-changing data such as price and availability. In practice that means many competitor product pages expose a price in the plain HTML response, so a plain HTTP fetch plus a JSON-LD parser covers a large share of targets before you need a browser.

This extractor reads Product JSON-LD, follows @graph, handles a single offer or a list, falls back to a price specification when there is no top-level price, and captures the strikethrough price separately. It skips AggregateOffer on purpose, because a price range is not one seller's price.

import json
from decimal import Decimal, InvalidOperation

from bs4 import BeautifulSoup


def to_decimal(value):
    """schema.org prices use a dot decimal and no thousands separators."""
    try:
        return Decimal(str(value).strip())
    except (InvalidOperation, AttributeError):
        return None


def iter_nodes(obj):
    """Yield every dict in a JSON-LD payload, including @graph members."""
    if isinstance(obj, list):
        for item in obj:
            yield from iter_nodes(item)
    elif isinstance(obj, dict):
        if "@graph" in obj:
            yield from iter_nodes(obj["@graph"])
        yield obj


def has_type(node, name):
    t = node.get("@type")
    return name in (t if isinstance(t, list) else [t])


def extract_offers(html):
    """Return one record per Offer found in Product JSON-LD on the page."""
    soup = BeautifulSoup(html, "html.parser")
    records = []
    for tag in soup.find_all("script", type="application/ld+json"):
        try:
            data = json.loads(tag.string or "")
        except json.JSONDecodeError:
            continue
        for node in iter_nodes(data):
            if not has_type(node, "Product"):
                continue
            offers = node.get("offers") or []
            if isinstance(offers, dict):
                offers = [offers]
            for offer in offers:
                if has_type(offer, "AggregateOffer"):
                    continue  # a price range, not one seller's price
                specs = offer.get("priceSpecification") or []
                if isinstance(specs, dict):
                    specs = [specs]
                price, currency = offer.get("price"), offer.get("priceCurrency")
                if price is None:  # active price may sit inside a spec
                    active = next(
                        (s for s in specs
                         if "priceType" not in s and "validForMemberTier" not in s),
                        None,
                    )
                    if active:
                        price, currency = active.get("price"), active.get("priceCurrency")
                strike = next(
                    (s for s in specs
                     if "StrikethroughPrice" in str(s.get("priceType", ""))),
                    None,
                )
                records.append({
                    "gtin": node.get("gtin13") or node.get("gtin14")
                            or node.get("gtin12") or node.get("gtin"),
                    "mpn": node.get("mpn"),
                    "sku": node.get("sku"),
                    "price": to_decimal(price),
                    "currency": currency,
                    "list_price": to_decimal(strike.get("price")) if strike else None,
                    "availability": str(offer.get("availability", "")).rsplit("/", 1)[-1],
                })
    return records

What it does not cover: microdata, product variants (variant pages can point to a group through isVariantOf), and sites with no markup at all. Those fall through to tier 2 or 3. Keep one extractor per tier behind the same output shape, and record which one produced each row so a drop in quality is traceable.

How to Match Competitor Products to Your SKUs

Matching decides whether a price comparison means anything. A cheaper price on the wrong product is worse than no data. Rank match methods by confidence and let confidence control what automation is allowed to do.

Match method Confidence Allowed use
Exact GTIN, EAN or UPC High Alerts and automated repricing
MPN plus brand Medium to high Alerts; automate after a spot check
Fuzzy title plus attributes Low Human review queue only

Traps that survive a correct identifier match:

  • Pack size. A 2-pack against a single unit. Normalise to unit price where the page gives a reference quantity.
  • Condition. New against refurbished or used; check itemCondition when present.
  • Variants. Size or colour variants can carry different prices under one product name.
  • Bundles. A kit that includes your item plus accessories.

Store the match method on every mapping, as in the schema below. When a match is later disproved, you can find and review every row that used the same method.

How Often to Scrape and How to Store Price History

Scrape frequency should follow how fast prices move, not a single global cron. Treat these as starting points and tune them from observed change rates per competitor.

SKU tier Example Starting frequency
Price-critical Best sellers, products with many direct rivals Daily, or several times a day
Standard Mid-volume catalog Daily
Long tail Slow movers, unique items Weekly

Store an append-only observation for every successful fetch, including unchanged prices, so you can distinguish "no change" from "scraper did not run". Keep the extractor and HTTP status on the row. A minimal PostgreSQL schema:

CREATE TABLE competitor_offer (
  id            BIGSERIAL PRIMARY KEY,
  competitor_id INT  NOT NULL,
  url           TEXT NOT NULL,
  our_sku       TEXT,
  match_method  TEXT,          -- gtin | mpn_brand | manual
  UNIQUE (competitor_id, url)
);

CREATE TABLE price_observation (
  offer_id     BIGINT      NOT NULL REFERENCES competitor_offer(id),
  observed_at  TIMESTAMPTZ NOT NULL,
  price        NUMERIC(12,2),
  list_price   NUMERIC(12,2),
  currency     CHAR(3),
  availability TEXT,
  shipping     NUMERIC(12,2),
  extractor    TEXT,           -- jsonld | css | rendered
  http_status  SMALLINT,
  PRIMARY KEY (offer_id, observed_at)
);

Derive "price changed" events from consecutive observations rather than overwriting a current-price column. History is what lets you answer questions such as how often a rival runs promotions, or whether a drop was a one-day sale.

Why Scrapers Break and How to Detect It

The expensive failure is not a crash. It is a scraper that keeps running and quietly writes wrong prices. Monitor the scraper as carefully as you monitor the prices.

Symptom Likely cause Detection
Many null prices from one competitor Layout or markup change Null-rate per competitor per run, with an alert threshold
Price jumps tenfold or drops to zero Parsing error, "from" price, wrong element Sanity band against the last observation; hold outliers for review
Spike of 403 or 429 responses Rate limiting or blocking Status-code histogram per host; slow down and back off
Same price for every product Consent page or interstitial fetched instead of the product Page fingerprint check; reject known interstitial markers
Prices in an unexpected currency Site serves region-based pricing Pin the locale or per-currency URL; validate currency per row
Extracted price differs from the visible one Stale or tax-exclusive markup Scheduled sample audit against the rendered page

When a competitor starts blocking you, slowing down and reducing frequency is the right response. Escalating to evasion tactics raises both the legal and the maintenance cost.

Turning Price Data into Repricing Decisions

Raw competitor prices should feed rules, not replace judgment. A few guardrails that keep automation from hurting margin:

  • Compare landed price (item plus shipping), not the sticker price alone.
  • Set a floor from cost plus minimum margin, and respect any minimum advertised price you are bound by.
  • Ignore out-of-stock competitors and low-confidence matches.
  • Require human approval for moves beyond a set percentage.
  • Log every automated change with the observations that triggered it.

Start with alerts only. Once the match rate and failure alerts have been quiet for a few weeks, let GTIN-matched, in-stock products reprice automatically inside the floor, and keep everything else in review.

If you want a scraper built around your own competitors, catalog and pricing rules, send the sites and SKUs that matter through Netalith's free quote form and we will scope it with you.

FAQ

Frequently asked questions

Is it legal to scrape competitor prices?

It depends on the country, the site's terms, what you access and how. Prices on public pages are commonly monitored, but terms of service, robots.txt, logins and bot protection all matter, and robots.txt is not access authorization. Do not scrape behind a login or bypass protections, and get legal advice before scraping at scale.

How often should I scrape competitor prices?

Match frequency to how fast prices move. A common starting point is daily or more often for price-critical SKUs and weekly for the long tail, then adjust from each competitor's observed change rate.

Should I build a custom scraper or buy a price monitoring tool?

Buy when you track a few mainstream retailers and want data quickly. Build when you monitor many niche or regional sites, need your own product matching, or must feed prices straight into your own database and repricing rules. Many sellers combine both.

How do I match competitor products to my own SKUs?

Prefer exact GTIN, EAN or UPC, then MPN plus brand. Use fuzzy title matching only to suggest candidates for human review. Watch for pack sizes, condition and variants, and store how each match was made.

What if a competitor blocks my scraper?

Slow down, lower the frequency and check that you follow their robots.txt. If the site still excludes you, use another legitimate source such as an official API, a licensed feed or manual sampling rather than escalating to evasion.

Stay updated with Netalith

Get coding resources, product updates, and special offers directly in your inbox.