Ecommerce Competitor Price Monitoring: How to Build a Custom Scraper That Lasts
Ecommerce competitor price monitoring with a custom scraper: which prices to capture, how to extract and match them, store history, and respect robots.txt.
Long Nguyen
Fullstack Developer · AI Engineer · Researcher
What a Competitor Price Monitoring Scraper Has to Do
A custom scraper for ecommerce competitor price monitoring is a pipeline, not a script. It fetches competitor product pages, extracts price and availability, matches each page to one of your SKUs, stores every observation, and tells you when something changed or broke.
Fetching and parsing is the easy part. Whether the data can be trusted depends on matching, history and failure detection:
- Fetch product URLs politely, on a schedule.
- Extract price, currency, list price, availability and shipping.
- Match each competitor URL to your own SKU.
- Store every observation, not just the latest price.
- Alert on meaningful price moves and on scraper failures.
The rest of this guide follows that order.
Custom Scraper vs Price Monitoring Tool: When Building Wins
A hosted monitoring tool is usually the faster start. A custom scraper wins when your problem does not fit a tool's assumptions.
| Factor | Custom scraper | Hosted monitoring tool |
|---|---|---|
| Competitor sites | Many niche, regional or oddly built sites | A few mainstream retailers the tool already covers |
| Product matching | Your rules: GTIN, MPN, internal part numbers, bundles | The tool's matching, with limited control |
| Integration | Straight into your database, catalog or repricing rules | Export or API within the tool's limits |
| History and ownership | Full history in your own database | Depends on plan and retention |
| Cost shape | Engineering time up front, then maintenance | Recurring subscription, little engineering |
| Maintenance | Yours: layout changes, blocks, monitoring | The vendor's |
| Time to first data | Longer | Shorter |
A sensible split for many sellers: use a tool for the handful of big competitors it handles well, and a scraper for the long tail of niche sites and for anything that must feed your own pricing logic.
If you decide to build and would rather not run the crawler yourself, Netalith's custom software service covers data crawling and processing, including price collection pipelines like the one described here.
Which Prices and Fields to Capture
Do not store one number called "price". Google's merchant listing documentation distinguishes three kinds of price: the active price, a strikethrough price shown during a sale, and a member price for a loyalty program. It encodes them in Offer and UnitPriceSpecification markup, using priceType for the strikethrough price and validForMemberTier for the member price. Mixing them is a reliable way to produce a wrong comparison.
| Field | Why it matters | Typical source |
|---|---|---|
| Active price | What a shopper pays right now | price, or a price specification without priceType |
| Strikethrough price | Separates a real cut from a permanent "was" price | Price specification with priceType set to StrikethroughPrice |
| Member price | A loyalty price is not the public price | validForMemberTier |
| Currency | Never compare raw numbers across currencies | priceCurrency |
| Availability | An out-of-stock rival is not a real price threat | availability |
| Shipping cost and delivery time | Needed to compare landed price | shippingDetails, or visible page text |
| Sale window | Tells you when a promotion ends | validFrom, priceValidUntil, validThrough |
| Unit price | Compares per-100 ml or per-kilo, not per-pack | referenceQuantity |
Two practical notes. When a store sells in several currencies, Google's guidance is one URL per currency, so key your observations on URL plus currency. And markup can lag the visible page or leave out tax, so audit a sample of extracted prices against what a shopper actually sees.
How to Extract Prices: Structured Data First, HTML Second
Try extraction methods from cheapest and most stable to most expensive:
- Product structured data (JSON-LD or microdata) in the page source. Stable, cheap, and often carries GTIN and MPN too.
- Site-specific CSS or XPath selectors for sites without usable markup.
- A rendered page in a headless browser, only when the price is injected by JavaScript.
Google recommends putting Product markup in the initial HTML, and warns that JavaScript-generated markup can make Shopping crawls less frequent and less reliable for fast-changing data such as price and availability. In practice that means many competitor product pages expose a price in the plain HTML response, so a plain HTTP fetch plus a JSON-LD parser covers a large share of targets before you need a browser.
This extractor reads Product JSON-LD, follows @graph, handles a single offer or a list, falls back to a price specification when there is no top-level price, and captures the strikethrough price separately. It skips AggregateOffer on purpose, because a price range is not one seller's price.
import json
from decimal import Decimal, InvalidOperation
from bs4 import BeautifulSoup
def to_decimal(value):
"""schema.org prices use a dot decimal and no thousands separators."""
try:
return Decimal(str(value).strip())
except (InvalidOperation, AttributeError):
return None
def iter_nodes(obj):
"""Yield every dict in a JSON-LD payload, including @graph members."""
if isinstance(obj, list):
for item in obj:
yield from iter_nodes(item)
elif isinstance(obj, dict):
if "@graph" in obj:
yield from iter_nodes(obj["@graph"])
yield obj
def has_type(node, name):
t = node.get("@type")
return name in (t if isinstance(t, list) else [t])
def extract_offers(html):
"""Return one record per Offer found in Product JSON-LD on the page."""
soup = BeautifulSoup(html, "html.parser")
records = []
for tag in soup.find_all("script", type="application/ld+json"):
try:
data = json.loads(tag.string or "")
except json.JSONDecodeError:
continue
for node in iter_nodes(data):
if not has_type(node, "Product"):
continue
offers = node.get("offers") or []
if isinstance(offers, dict):
offers = [offers]
for offer in offers:
if has_type(offer, "AggregateOffer"):
continue # a price range, not one seller's price
specs = offer.get("priceSpecification") or []
if isinstance(specs, dict):
specs = [specs]
price, currency = offer.get("price"), offer.get("priceCurrency")
if price is None: # active price may sit inside a spec
active = next(
(s for s in specs
if "priceType" not in s and "validForMemberTier" not in s),
None,
)
if active:
price, currency = active.get("price"), active.get("priceCurrency")
strike = next(
(s for s in specs
if "StrikethroughPrice" in str(s.get("priceType", ""))),
None,
)
records.append({
"gtin": node.get("gtin13") or node.get("gtin14")
or node.get("gtin12") or node.get("gtin"),
"mpn": node.get("mpn"),
"sku": node.get("sku"),
"price": to_decimal(price),
"currency": currency,
"list_price": to_decimal(strike.get("price")) if strike else None,
"availability": str(offer.get("availability", "")).rsplit("/", 1)[-1],
})
return records
What it does not cover: microdata, product variants (variant pages can point to a group through isVariantOf), and sites with no markup at all. Those fall through to tier 2 or 3. Keep one extractor per tier behind the same output shape, and record which one produced each row so a drop in quality is traceable.
How to Match Competitor Products to Your SKUs
Matching decides whether a price comparison means anything. A cheaper price on the wrong product is worse than no data. Rank match methods by confidence and let confidence control what automation is allowed to do.
| Match method | Confidence | Allowed use |
|---|---|---|
| Exact GTIN, EAN or UPC | High | Alerts and automated repricing |
| MPN plus brand | Medium to high | Alerts; automate after a spot check |
| Fuzzy title plus attributes | Low | Human review queue only |
Traps that survive a correct identifier match:
- Pack size. A 2-pack against a single unit. Normalise to unit price where the page gives a reference quantity.
- Condition. New against refurbished or used; check
itemConditionwhen present. - Variants. Size or colour variants can carry different prices under one product name.
- Bundles. A kit that includes your item plus accessories.
Store the match method on every mapping, as in the schema below. When a match is later disproved, you can find and review every row that used the same method.
How Often to Scrape and How to Store Price History
Scrape frequency should follow how fast prices move, not a single global cron. Treat these as starting points and tune them from observed change rates per competitor.
| SKU tier | Example | Starting frequency |
|---|---|---|
| Price-critical | Best sellers, products with many direct rivals | Daily, or several times a day |
| Standard | Mid-volume catalog | Daily |
| Long tail | Slow movers, unique items | Weekly |
Store an append-only observation for every successful fetch, including unchanged prices, so you can distinguish "no change" from "scraper did not run". Keep the extractor and HTTP status on the row. A minimal PostgreSQL schema:
CREATE TABLE competitor_offer (
id BIGSERIAL PRIMARY KEY,
competitor_id INT NOT NULL,
url TEXT NOT NULL,
our_sku TEXT,
match_method TEXT, -- gtin | mpn_brand | manual
UNIQUE (competitor_id, url)
);
CREATE TABLE price_observation (
offer_id BIGINT NOT NULL REFERENCES competitor_offer(id),
observed_at TIMESTAMPTZ NOT NULL,
price NUMERIC(12,2),
list_price NUMERIC(12,2),
currency CHAR(3),
availability TEXT,
shipping NUMERIC(12,2),
extractor TEXT, -- jsonld | css | rendered
http_status SMALLINT,
PRIMARY KEY (offer_id, observed_at)
);
Derive "price changed" events from consecutive observations rather than overwriting a current-price column. History is what lets you answer questions such as how often a rival runs promotions, or whether a drop was a one-day sale.
robots.txt, Terms and Legal Limits
Build the crawler to follow RFC 9309, the Robots Exclusion Protocol, an IETF standards-track document from September 2022. The parts that matter for a price scraper:
- The standard states that its rules are not a form of access authorization. Honoring robots.txt is a courtesy and a baseline, not a legal clearance.
- If robots.txt returns a 4xx status, the crawler may access any resource. If it is unreachable because of a 5xx or a network error, the crawler must assume complete disallow.
- Cache robots.txt, but not for more than 24 hours unless it is unreachable.
- The most specific matching rule wins, and an equivalent allow beats a disallow.
- Use a product token that appears in your User-Agent and describes your crawler's purpose, ideally with a contact URL.
- The standard defines only allow and disallow rules, so request pacing is your own responsibility: limit concurrency per host and back off on errors.
Beyond robots.txt, three rules keep the project defensible. Read each competitor's terms of service. Do not scrape pages behind a login, and do not work around CAPTCHAs or bot protection. If a competitor's robots.txt or terms exclude product pages, treat that as a no and look for another legitimate source, such as an official marketplace API, a feed you are licensed to use, or manual sampling.
Legal treatment of scraping differs by country and by what is scraped, so this is not legal advice. Take advice before scraping at scale or in a new jurisdiction. Separately, monitoring public prices is different from coordinating prices with competitors; keep the data as an input to your own independent pricing decisions.
Why Scrapers Break and How to Detect It
The expensive failure is not a crash. It is a scraper that keeps running and quietly writes wrong prices. Monitor the scraper as carefully as you monitor the prices.
| Symptom | Likely cause | Detection |
|---|---|---|
| Many null prices from one competitor | Layout or markup change | Null-rate per competitor per run, with an alert threshold |
| Price jumps tenfold or drops to zero | Parsing error, "from" price, wrong element | Sanity band against the last observation; hold outliers for review |
| Spike of 403 or 429 responses | Rate limiting or blocking | Status-code histogram per host; slow down and back off |
| Same price for every product | Consent page or interstitial fetched instead of the product | Page fingerprint check; reject known interstitial markers |
| Prices in an unexpected currency | Site serves region-based pricing | Pin the locale or per-currency URL; validate currency per row |
| Extracted price differs from the visible one | Stale or tax-exclusive markup | Scheduled sample audit against the rendered page |
When a competitor starts blocking you, slowing down and reducing frequency is the right response. Escalating to evasion tactics raises both the legal and the maintenance cost.
Turning Price Data into Repricing Decisions
Raw competitor prices should feed rules, not replace judgment. A few guardrails that keep automation from hurting margin:
- Compare landed price (item plus shipping), not the sticker price alone.
- Set a floor from cost plus minimum margin, and respect any minimum advertised price you are bound by.
- Ignore out-of-stock competitors and low-confidence matches.
- Require human approval for moves beyond a set percentage.
- Log every automated change with the observations that triggered it.
Start with alerts only. Once the match rate and failure alerts have been quiet for a few weeks, let GTIN-matched, in-stock products reprice automatically inside the floor, and keep everything else in review.
If you want a scraper built around your own competitors, catalog and pricing rules, send the sites and SKUs that matter through Netalith's free quote form and we will scope it with you.
FAQ
Frequently asked questions
Is it legal to scrape competitor prices?
It depends on the country, the site's terms, what you access and how. Prices on public pages are commonly monitored, but terms of service, robots.txt, logins and bot protection all matter, and robots.txt is not access authorization. Do not scrape behind a login or bypass protections, and get legal advice before scraping at scale.
How often should I scrape competitor prices?
Match frequency to how fast prices move. A common starting point is daily or more often for price-critical SKUs and weekly for the long tail, then adjust from each competitor's observed change rate.
Should I build a custom scraper or buy a price monitoring tool?
Buy when you track a few mainstream retailers and want data quickly. Build when you monitor many niche or regional sites, need your own product matching, or must feed prices straight into your own database and repricing rules. Many sellers combine both.
How do I match competitor products to my own SKUs?
Prefer exact GTIN, EAN or UPC, then MPN plus brand. Use fuzzy title matching only to suggest candidates for human review. Watch for pack sizes, condition and variants, and store how each match was made.
What if a competitor blocks my scraper?
Slow down, lower the frequency and check that you follow their robots.txt. If the site still excludes you, use another legitimate source such as an official API, a licensed feed or manual sampling rather than escalating to evasion.