UK grocery is the most mechanically complex retail market in Europe to collect data from, and the reason is that no two of the major retailers price the same way. Tesco and Sainsbury's run loyalty price layers. ASDA runs a cashpot that is not a discount at all. Morrisons sells a large share of its range by weight. Aldi and Lidl are almost entirely own-label with no shared barcodes, and much of Lidl's promotional activity sits behind an app. A schema built for one of them will silently produce wrong numbers on the others. This guide maps the landscape, compares the pricing mechanics side by side, sets out the matching methodology that cross-retailer comparison actually requires, and covers UK compliance.
The UK grocery market is dominated by four large supermarket groups — Tesco, Sainsbury's, ASDA and Morrisons — alongside two discounters, Aldi and Lidl, which have taken substantial share over the past decade — in the 12 weeks to 9 August 2026, Worldpanel by Numerator (formerly Kantar Worldpanel) put the six at Tesco 27.8%, Sainsbury's 15.2%, Asda 11.5%, Aldi 10.7%, Lidl 8.8% and Morrisons 8.5%, with Aldi and Lidl together now accounting for close to 19.5% of the market. Beyond them sit Waitrose, Co-op, Iceland, M&S Food and Ocado, each serving distinct positions.
The commercial demand for this data comes from five directions.
What unites them is that they all need the same thing: accurate, comparable, timestamped price data across multiple retailers. And that turns out to be much harder than it sounds.
This is the section that matters most, and it is the thing most guides skip.
There is no such thing as "a UK supermarket price schema". Each retailer runs a different promotional architecture, and those differences are not cosmetic — they change what a price field means. Here is the landscape in one view.
| Retailer | Primary mechanic | Is it a discount? | Key schema requirement |
|---|---|---|---|
| Tesco | Clubcard Price | Yes — lower price for members | Dual price fields, offer conditions |
| Sainsbury's | Nectar Price + Aldi Price Match | Yes — both lower the price | Three overlapping price layers |
| ASDA | Rollback + Rewards cashpot | Rollback yes, Rewards no | Rewards must sit outside price |
| Morrisons | More Card + Price Lock | More Card yes, Price Lock no | Lock is duration, not discount |
| Aldi | Simple pricing, Super 6 | Yes — straightforward reductions | Ephemeral Specialbuys lifecycle |
| Lidl | Weekly offers + Lidl Plus app | Public yes, app-gated off-limits | Source provenance per price point |
Read down the third column. Three of the six run a headline mechanic that is not a price reduction at all, and each fails differently if you model it as one.
A single "effective price" column applied across all six retailers will be wrong on at least three of them, in different directions, by different amounts. That is why cross-retailer grocery datasets so often fail quietly: nothing errors, the file validates, and the numbers are simply not comparable.
Beyond the per-retailer mechanics, five issues recur across the whole market. Every one of them is a source of silently wrong data.
UK online grocery pricing and availability resolve in the context of a delivery postcode or a selected store. Change the context and you can change availability, offers and — at retailers operating convenience formats — the price itself. ASDA Express and Morrisons Daily generally price above their large-store equivalents.
A capture without a controlled, recorded location context is not reproducible. You cannot compare Tuesday's file to Wednesday's and trust the delta, because you do not know whether the price moved or the context did.
The rule: every capture is pinned to an explicit location and store format, and both ship as fields on every row, stamped at collection time rather than reconstructed later. For national benchmarking, fix one reference context and hold it constant for the life of the dataset. For regional work, run a defined panel in parallel. Never blend formats into one price series.
Clubcard Prices, Nectar Prices and More Card prices are genuine price reductions available to members. ASDA Rewards is not. Lidl Plus is not public. Treating "loyalty" as a single concept across retailers produces a dataset where the same column means four different things.
The rule: model pricing as an array of typed offer objects rather than a fixed set of columns, then flatten to columns at output. A new promotional mechanic next year becomes a new entry, not a schema migration and a broken downstream report.
Morrisons in particular sells a large share of its range by weight — Market Street counter lines, loose produce, variable-weight packs. A whole chicken at £6.50 per kilogram recorded into a price column that also holds a £1.20 tin of beans corrupts every category average by a factor that varies per product.
The rule: pricing_basis is a required field with no default. If it cannot be resolved for a row, that row fails validation rather than defaulting to per-item. Normalise everything to a comparable unit price at ingestion — that is the only field that compares across fixed and variable weight lines, and across retailers.
Aldi Specialbuys and Middle of Lidl are time-limited non-food ranges that drop on a schedule, sell out, and disappear. A product that appears and sells out between two daily captures never existed as far as your dataset is concerned — and there is no archive to backfill from.
The rule: capture cadence follows the drop schedule, not the clock. Sell-out timing is often more commercially valuable than price, and it exists only if you sampled densely enough to observe both ends of the transition. Ship the observation gap alongside any derived duration as an uncertainty band.
This is the problem that defines cross-retailer work, and it gets its own section below.
Comparing prices across UK grocers sounds like a lookup problem. It is not. It is a matching problem, and the difficulty scales with own-label share.
Branded lines match cleanly. A branded product carries the same EAN at Tesco, Sainsbury's, ASDA and Morrisons. Match on identifier, compare prices, done. This is the easy 30-40% of a typical basket.
Own-label lines do not match at all. A Tesco own-label bean has no equivalent EAN at Sainsbury's. There is no identical item — there are three different products serving the same need at three different quality positions.
Discounter comparison is hardest, because both sides are almost entirely own-label. An Aldi-versus-Lidl comparison has no shared identifier anywhere. Every single matched pair is an attribute-based judgement.
If your output supports public "cheaper than" claims, the matching methodology carries legal weight. Misleading comparative advertising is regulated in the UK. A comparison whose methodology is unstated and whose confidence is unexposed is not defensible when a competitor's legal team challenges it — and in UK grocery, they do.
Any serious comparison dataset ships its matching logic alongside its prices. If a vendor will not show you theirs, that is the answer to your question.
The practical approach is a common core plus retailer-specific extensions. Every row carries the core; retailer-specific fields populate where they apply.
Core identity: product_id, product_url, product_name, brand, is_own_label, own_label_tier, pack_size, gtin_ean, category_path, image_urls, retailer.
Core pricing: price, currency, pricing_basis, unit_price, unit_of_measure, was_price, offers[] (typed array), effective_price (derived, till-price only).
Core context: availability_status, delivery_postcode, store_id, store_format, fulfilment_type, price_source, captured_at.
Retailer extensions: clubcard_price (Tesco); nectar_price, aldi_price_match (Sainsbury's); is_rollback, rewards_offer, rewards_conditions (ASDA); more_card_price, is_price_lock, price_lock_until, price_per_kg, estimated_weight_kg (Morrisons); is_specialbuy, sold_out_at, availability_duration_hours (Aldi); is_middle_of_lidl, observation_gap_minutes (Lidl).
Three design rules make this work.
Uniform daily capture across everything is the most common configuration and usually the wrong one — it overspends on stable categories and underspends where it matters.
| Category type | Recommended cadence | Why |
|---|---|---|
| Ambient packaged goods | Daily | Stable; weekly promo cycles |
| Fresh and counter lines | Twice daily | Availability shifts through the day |
| Promotional monitoring | Daily minimum | UK promo cycles turn over weekly |
| Ephemeral ranges (Specialbuys, Middle of Lidl) | Drop-aligned, 15-min intervals at launch | Sell-out timing cannot be backfilled |
| Inflation research | Weekly sufficient | Depth of history beats frequency |
| Availability alerting | Intraday | Stock state is the point |
The asymmetry worth internalising: some fields cannot be created later. Rollback duration, Specialbuy sell-out speed, promotional window length — these are derived from observed state transitions. A client who starts with weekly capture and later wants duration analysis cannot backfill it. The data was never collected.
That makes the cost of starting late genuinely irrecoverable, which is not true of most data projects.
Front ends change. Parsers break. The dangerous failures are the ones that do not throw errors.
Boolean fields are the most dangerous fields in a retail scraper, because a broken parser and a real "no" look identical. When ASDA's Rollback selector stops matching, is_rollback writes false across the catalogue. Nothing errors. Three weeks later someone asks why ASDA appears to have stopped running promotions, and you have three weeks of corrupt history that cannot be recovered.
A validation layer that runs on every batch:
UK enterprise buyers raise this during procurement. It is the main commercial risk in the category.
Public data only. Collect what any visitor can see without authenticating. No logged-in areas, no account data, no basket data. This boundary does real work at Lidl, where a substantial share of promotional activity sits behind an authenticated app — that data is out of scope regardless of technical feasibility.
No personal data. Prices are not personal data. Customer reviews may contain reviewer names or identifiable content — if you collect reviews, UK GDPR applies and you need a lawful basis, retention policy and minimisation. For most price monitoring, collect review counts and averages only, never review text or author identity. Personalised offers are a different category again, because they are derived from an identified individual's behaviour.
Database rights. The UK retains a sui generis database right, separate from copyright, protecting substantial investment in obtaining, verifying or presenting database contents. Extracting a substantial part can infringe it. The defensible position is factual price monitoring for analysis and comparison — not republishing a retailer's catalogue as your own product.
Comparative claims. Misleading comparative advertising is regulated. Where output supports public price comparisons, matching methodology and confidence scores are part of the compliance posture, not just the analytics.
Rate limiting as a legal posture. Conduct that impairs a service is where scraping disputes escalate. Conservative volumes are a risk control — with particular care around high-traffic moments like discounter drop mornings.
Terms of service. Site terms are contractual and enforceability against non-account-holders varies. Note that this variability is precisely why account-gated content sits on a different and much less favourable footing.
Not legal advice. Take advice from a qualified UK solicitor for your specific programme.
The honest dividing lines, by scope.
Build is reasonable when: you need one or two ambient categories, from one or two retailers, refreshed weekly, and can tolerate gaps when a site changes. The first version of a UK grocery scraper is genuinely not hard.
Buy makes sense when: your scope includes any of the following — multi-retailer comparison requiring a matching layer; fresh and variable-weight categories; ephemeral ranges with drop-aligned capture; multi-format or regional panels; or an SLA someone else is accountable for.
The calculation that usually decides it is not the build cost. It is the three-year maintenance cost, and the cost of decisions made on wrong data before anyone noticed a break. Model both.
There is also a governance dimension. An in-house team under pressure to "just get the Lidl Plus data" faces a temptation that an external partner with a contractual scope boundary does not. Having the boundary written into a contract and into the validation layer is worth something on its own.
Use these in evaluation. The answers separate operators from resellers.
Question 7 is the fastest filter. A vendor who claims complete Lidl Plus coverage is either describing something narrower than it sounds or doing something that creates exposure for you. Either way, you have learned what you needed to know.
Tesco, Sainsbury's, ASDA, Morrisons, Aldi and Lidl are the core six. Waitrose, Co-op, Iceland, M&S Food and Ocado extend coverage. Each needs retailer-specific handling — a schema built for one will produce wrong numbers on the others.
Daily covers most promotional monitoring, since UK promo cycles typically turn over weekly. Fresh and availability tracking benefit from twice-daily. Ephemeral ranges need drop-aligned intraday capture. Inflation research is fine weekly, because depth of history matters more than frequency.
Yes, by running a defined postcode panel in parallel with the location and store format stamped on every row. Because UK grocery pricing resolves against a delivery location, this is the only reliable way to produce regional comparisons.
For Tesco Clubcard, Sainsbury's Nectar and Morrisons More Card, yes — these are displayed publicly on product pages so shoppers see the saving before signing in. Lidl Plus is different: it is app-gated and frequently personalised, which places it outside public data.
Through attribute-based matching on normalised unit price, own-label tier equivalence, pack size and published attributes, with a confidence score on every pair. Barcode matching does not work, because own-label products are exclusive to each retailer by definition.
Collecting publicly displayed factual pricing for analysis is a widely practised commercial activity. The risk areas are personal data under UK GDPR, database rights, contractual site terms, conduct that impairs a service, and — where output supports public comparison claims — comparative advertising rules. Take legal advice for your specific programme.
Applying one schema across all six retailers. It produces files that validate cleanly and are wrong in different directions on different retailers, which is far more damaging than an obvious failure because nobody catches it until a decision has already been made.
Each of these goes deep on one retailer's mechanics, schema and failure modes:
Actowiz Solutions delivers UK grocery datasets across all six major retailers, with retailer-specific mechanics modelled correctly, location and store-format context on every row, attribute-based cross-retailer matching with exposed confidence scores, validated schemas and scheduled delivery to S3, SFTP, BigQuery or API.
You can also reach us for all your mobile app scraping, data collection, web scraping , and instant data scraper service requirements!
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Explore the 2026 guide to UK supermarket price and product data, covering product catalogs, pricing, promotions, availability, and competitive insights.
Multi-Channel Marketplace Inventory Scraping API helps brands monitor product stock, availability, and inventory changes across Amazon, Flipkart, and Myntra.
Sephora & Trendyol Arabic Market Data Report 2026 delivers UAE e-commerce intelligence on products, pricing, trends, and customer demand.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.