ASDA is one of the big four UK grocers — holding around 11.4% of UK grocery market share, per Kantar Worldpanel's most recently published 12-week reading, behind Tesco, Sainsbury's and Aldi and historically the one most associated with everyday low pricing rather than promotional depth. That positioning shapes the data in a way most teams do not anticipate.
Tesco and Sainsbury's run loyalty pricing: a second, lower price displayed on the page, available to members. It is a price, and you capture it as a price field. ASDA's primary promotional mechanic is Rollback — a temporary reduction from a previous price, visible to everyone, with no loyalty gate. Structurally that is simpler. It behaves like a standard was/now reduction with a branded label.
But ASDA also runs ASDA Rewards, and this is where data teams get it wrong. Rewards does not lower the price at the till. It credits pounds into a cashpot that the shopper redeems later, typically triggered by buying specific "Star Product" lines or completing missions. The shopper still pays the shelf price today.
This distinction matters enormously for analysis. If you model a £2 Rewards offer as a £2 discount, you will report ASDA's effective prices as far lower than they actually are, and any basket comparison against Tesco or Sainsbury's becomes meaningless. Rewards is a loyalty incentive with a monetary value, captured as a separate field with its own type — never folded into effective_price.
We have seen more than one in-house dataset get this wrong and produce a competitor basket analysis that was off by several percent across an entire category. The error is invisible until someone checks a receipt.
Here is the field schema we run on production ASDA feeds.
| Field | Description | Example |
|---|---|---|
| product_id | ASDA internal SKU identifier | 910001234567 |
| product_url | Canonical product page URL | https://groceries.asda.com/product/... |
| product_name | Full product title as displayed | ASDA British Semi Skimmed Milk 2.27L |
| brand | Brand name, parsed or from structured data | ASDA |
| own_label_tier | Own-brand tier where applicable | Just Essentials / ASDA / Extra Special |
| pack_size | Size, weight or volume as listed | 2.27L |
| gtin_ean | Barcode identifier, where published | 05051413000000 |
| category_path | Full breadcrumb hierarchy | Fresh Food > Milk, Butter & Eggs > Fresh Milk |
| image_urls | Array of product image URLs | ["https://ui.assets-asda.com/..."] |
ASDA's own-label architecture runs from Just Essentials at the value end, through the core ASDA brand, to Extra Special at the premium end. Capture the tier explicitly. Just Essentials in particular is a heavily tracked range — it is the line most directly aimed at Aldi and Lidl, and its price movements are a leading indicator of discounter competitive pressure across the whole market.
| Field | Description | Example |
|---|---|---|
| price | Current shelf price | 1.65 |
| currency | ISO currency code | GBP |
| unit_price | Price per standard unit | 0.73 |
| unit_of_measure | Basis for the unit price | per litre |
| was_price | Previous price where a reduction is shown | 1.95 |
| is_rollback | Boolean — is this product on Rollback | true |
| rollback_end_date | End date of the Rollback where published | 2026-04-02 |
| promo_type | Nature of the offer | rollback / multibuy / price_drop / rewards_star_product |
| promo_text | Raw offer text exactly as displayed | Rollback. Was £1.95 |
| rewards_offer | Cashpot value where the product is a Rewards line | 2.00 |
| rewards_offer_type | Nature of the Rewards mechanic | star_product / mission |
| rewards_conditions | Qualifying condition text | Buy 3, get £2 in your Cashpot |
| savings_vs_was | Derived — was price minus current price | 0.30 |
Note the deliberate separation. price and was_price describe what the shopper pays. rewards_offer describes value the shopper receives later, under conditions. These are never summed into a single effective price. If a downstream user wants a "total value" view, they can compute it themselves from the raw fields — but the raw fields must survive intact, or nobody can reconstruct what actually happened.
| Field | Description | Example |
|---|---|---|
| availability_status | Stock state at time of capture | in_stock / out_of_stock / unavailable |
| delivery_postcode | Postcode context used for this capture | LS11 5AD |
| store_id | Store identifier where resolvable | 4520 |
| store_format | Format of the store context | superstore / supermarket / express |
| fulfilment_type | Collection method for this capture | delivery / click_collect |
| rating_average | Average customer rating | 4.4 |
| review_count | Number of reviews | 938 |
| captured_at | UTC timestamp of the capture | 2026-03-04T06:55:31Z |
The store_format field is the one most competitors leave out entirely, and it is the reason their ASDA datasets do not reconcile. More on that below.
Product description, ingredients, allergen statement, nutrition panel per 100g, storage and usage instructions, country of origin, and dietary flags. Brands use these for content compliance — checking whether the listing ASDA runs matches the content the brand supplied. An allergen mismatch is a safety issue, not a merchandising one.
This is the ASDA-specific problem and the one that quietly ruins datasets.
UK convenience formats generally price above large-store formats — smaller stores carry higher cost-to-serve, and that flows through to shelf price. ASDA operates across superstores, supermarkets and ASDA Express convenience sites. The consequence for a data pipeline is that "the ASDA price" is not a single number. It depends which store context resolved the request.
A team that captures without controlling for this produces a dataset where the same SKU appears at different prices on different days for no visible reason, because the underlying store context shifted. The analysts then spend a month building rules to "smooth out the noise" — and in doing so they smooth out real price movements too.
The fix is architectural and non-negotiable:
If you take one thing from this article, take this: a UK grocery price dataset without an explicit, stored location and format context is not a dataset, it is a collection of unrelated observations.
Covered above, but it bears repeating because it is the most common modelling error on ASDA.
Concretely: a product at £3.00 with a "Buy 3, get £2 in your Cashpot" offer is not £2.33 per unit. The shopper pays £9.00 today and receives £2 of credit redeemable later, conditional on completing the purchase of three units and on redeeming the cashpot before expiry. Those are different things with different commercial meanings, and a brand's trade team needs them separated to have any sensible conversation about promotional funding.
Store rewards_offer, rewards_offer_type and rewards_conditions as their own fields. Leave effective_price reflecting only what is paid at the till.
Rollback is ASDA's headline promotional mechanic, and tracking Rollback duration is genuinely valuable — it tells you how long a competitor sustained a price position, which is exactly what a category manager wants to know before agreeing to match it.
The problem is that an end date is not always published on the page. Where it is absent, you cannot record the duration directly. You have to derive it from observation: the Rollback began on the first capture where is_rollback flipped to true, and ended on the first capture where it flipped back to false.
That derivation only works if you are capturing daily and storing history with continuity. It is a good example of why refresh frequency is not just about freshness — some fields simply cannot exist without a consistent historical series behind them. A client who starts with weekly capture and later wants Rollback duration analysis cannot backfill it. The data was never collected.
The ASDA online catalogue runs to tens of thousands of active grocery SKUs (trade press puts the core grocery range at roughly 25,000–30,000 SKUs, with ASDA's chair confirming in December 2025 an active simplification programme targeting a reduction toward around 24,000–25,000), and ASDA additionally sells general merchandise and clothing under the George brand. Depending on entry point, general merchandise lines can surface alongside grocery results.
If your scope is grocery price monitoring, you need explicit category-scope rules to exclude George lines, or your SKU counts will drift and your category-level averages will be polluted by products that are not food. Conversely, if you are tracking UK clothing or homeware pricing, George is a valuable dataset in its own right and deserves a separate schema — clothing needs size and colour variant fields that grocery does not.
Beyond that, the standard scale problems apply: category-tree discovery that re-walks the hierarchy rather than trusting a static seed list; delisting detection that records a disappearance as delisted rather than letting rows silently vanish; change reconciliation on a stable key; and pack-size change detection, which is the shrinkflation signal and only surfaces if pack_size and unit_price are stored historically.
An illustrative record showing output schema. Values are synthetic, shown to demonstrate field shape and types — they do not represent live ASDA pricing. Request a live sample for real current data.
{
"product_id": "910001234567",
"product_url": "https://groceries.asda.com/product/example-product",
"product_name": "Example Brand Baked Beans 415g",
"brand": "Example Brand",
"own_label_tier": null,
"pack_size": "415g",
"gtin_ean": "05051413000000",
"category_path": "Food Cupboard > Tinned Food > Baked Beans",
"price": 1.20,
"currency": "GBP",
"unit_price": 0.29,
"unit_of_measure": "per 100g",
"was_price": 1.50,
"is_rollback": true,
"rollback_end_date": "2026-04-02",
"promo_type": "rollback",
"promo_text": "Rollback. Was £1.50",
"rewards_offer": null,
"rewards_offer_type": null,
"rewards_conditions": null,
"savings_vs_was": 0.30,
"availability_status": "in_stock",
"delivery_postcode": "LS11 5AD",
"store_id": "4520",
"store_format": "superstore",
"fulfilment_type": "delivery",
"rating_average": 4.4,
"review_count": 938,
"captured_at": "2026-03-04T06:55:31Z"
}
Flattened to CSV, as category teams prefer it:
| product_id | product_name | tier | price | was | rollback | rewards | format | availability | captured_at |
|---|---|---|---|---|---|---|---|---|---|
| 910001234567 | Baked Beans 415g | — | 1.20 | 1.50 | true | — | superstore | in_stock | 2026-03-04 |
| 910001234568 | Semi Skimmed Milk 2.27L | ASDA | 1.65 | — | false | — | superstore | in_stock | 2026-03-04 |
| 910001234569 | Value Pasta 500g | Just Essentials | 0.45 | — | false | — | superstore | in_stock | 2026-03-04 |
| 910001234570 | Premium Cheddar 350g | Extra Special | 4.50 | — | false | £2 cashpot | superstore | out_of_stock | 2026-03-04 |
| 910001234568 | Semi Skimmed Milk 2.27L | ASDA | 1.85 | — | false | — | express | in_stock | 2026-03-04 |
Look at rows two and five. Same SKU, same day, two different prices — £1.65 in a superstore, £1.85 in an Express store. Neither is wrong. A dataset that does not carry store_format cannot tell you that, and will instead present it as a 12% price increase that never happened.
Row four shows the Rewards separation working: the price stays at £4.50, and the £2 cashpot sits in its own field where it cannot contaminate the price series.
Read https://groceries.asda.com/robots.txt and honour it. Confine collection to publicly accessible pages — no logged-in areas, no account data, no ASDA Rewards account information, no personal data. If a path is disallowed, it is out of scope. This boundary is what makes the operation defensible.
Extract from structured data where it exists rather than from CSS selectors. Many retail product pages publish Product schema in JSON-LD, which gives name, brand, identifiers, images and price in a machine-readable form that changes far less often than front-end class names.
A simplified, courteous fetch-and-parse pattern:
import json, time, requests
from bs4 import BeautifulSoup
HEADERS = {"User-Agent": "ActowizDataBot/1.0 (+https://actowizsolutions.com/bot)"}
DELAY_SECONDS = 3 # conservative; stay well inside courteous limits
def parse_product(url: str) -> dict | None:
resp = requests.get(url, headers=HEADERS, timeout=30)
resp.raise_for_status()
soup = BeautifulSoup(resp.text, "html.parser")
for tag in soup.find_all("script", type="application/ld+json"):
try:
data = json.loads(tag.string or "")
except json.JSONDecodeError:
continue
if isinstance(data, dict) and data.get("@type") == "Product":
offer = data.get("offers") or {}
return {
"product_name": data.get("name"),
"brand": (data.get("brand") or {}).get("name"),
"gtin_ean": data.get("gtin13"),
"price": offer.get("price"),
"currency": offer.get("priceCurrency"),
"availability_status": offer.get("availability"),
"product_url": url,
}
return None
def crawl(urls: list[str], store_context: dict) -> list[dict]:
out = []
for u in urls:
if (record := parse_product(u)):
record.update(store_context) # stamp every row with its context
out.append(record)
time.sleep(DELAY_SECONDS) # rate limiting is not optional
return out
The store_context parameter is the important line. Every row gets stamped with the postcode, store id and format it was captured under, at collection time — not reconstructed later from logs. Context that is not stamped on the row will eventually be lost.
What this snippet deliberately does not do: it makes no attempt to evade any protective measure, and it does not handle Rollback or Rewards parsing. Those sit in the promotional presentation layer, which requires page-specific logic — and that is exactly the part that needs ongoing maintenance as the front end changes.
Usually around month three, and usually silently.
ASDA ships a front-end change, the Rollback badge selector stops matching, and is_rollback starts writing false across the whole catalogue. No error is thrown — false is a perfectly valid boolean. The feed looks healthy. Three weeks later someone asks why ASDA appears to have stopped running Rollbacks entirely, and the answer is that you have three weeks of corrupt promotional history that cannot be recovered.
Boolean fields are the most dangerous fields in a retail scraper precisely because a broken parser and a real "no" look identical. Validation has to run on every batch:
CSV and Excel for category and merchandising teams working in spreadsheets. JSON or JSONL for engineering pipelines. Parquet where volume is high and query cost matters.
S3, Google Cloud Storage or Azure Blob; SFTP for established file-drop workflows; direct load into BigQuery, Snowflake or Redshift; or a REST endpoint for on-demand querying.
A full snapshot ships the entire catalogue state each run. A change log ships only what moved, with change type recorded (price_change, rollback_started, rollback_ended, rewards_offer_added, rewards_offer_removed, stock_change, new_listing, delisted). Most mature programmes take a weekly full snapshot for reconciliation plus a daily change log for alerting.
For competitive response work the file is not the deliverable, the alert is. A category manager wants a Slack message the morning a Rollback starts on a competing line in their category — not a 40,000-row CSV that tells them about it on Friday.
CPG and FMCG brands monitor their SKUs for price compliance, promotional execution, availability and content accuracy. ASDA's everyday-low-price positioning means unexpected Rollbacks on a brand's lines are a frequent source of channel conflict — the brand needs to know the day it starts, not at the quarterly review.
Competing grocers benchmark baskets like-for-like. ASDA is the key benchmark for value positioning among the big four, and Just Essentials specifically is the range most closely watched as a proxy for discounter pressure.
Discounters (Aldi, Lidl) and value retailers track Just Essentials pricing directly, because it is the range aimed at them.
Price comparison and cashback platforms need broad catalogue coverage refreshed frequently enough that displayed prices are not stale — and need store-format context, or their comparisons will be challenged.
Analysts, researchers and journalists track food inflation at SKU level and study shrinkflation by pairing pack_size with unit_price historically. ASDA's value ranges are frequently used as an inflation bellwether for lower-income baskets.
UK enterprise buyers ask about this during procurement. It is the main commercial risk in the category.
Build in-house if you need one or two categories, refresh weekly, have a data engineer with genuine spare capacity, and can tolerate gaps when the site changes.
Buy a managed feed if you need full-catalogue coverage, daily refresh, multi-store-format capture, multi-retailer comparison across ASDA, Tesco, Sainsbury's, Morrisons, Aldi and Lidl, a guaranteed schema, and an SLA.
ASDA tips the build-versus-buy calculation harder than the others because of the store-format requirement. Capturing a single national price series is a manageable in-house project. Capturing a controlled panel of store contexts, daily, with format stamped on every row and validation that catches a silent boolean failure — that is a standing engineering commitment, not a project with an end date.
Model it over three years, not three months, and include the cost of decisions made on wrong data before anyone noticed the break.
Rollback is a temporary reduction in the shelf price, visible to all shoppers, captured as a lower price with a was_price and an is_rollback flag. Rewards credits pounds into a cashpot for later redemption and does not reduce the price paid today — it is captured in separate rewards_offer fields and must never be folded into the effective price.
Because ASDA prices vary by store format. Convenience formats such as ASDA Express generally price above superstores. Any dataset without a recorded store format and location context will show these as unexplained price movements.
Yes, but only with daily capture and continuous history. End dates are not consistently published, so duration is derived by observing when is_rollback flips true and when it flips back. It cannot be backfilled later — if the daily series was not collected, the duration data does not exist.
ASDA has not historically offered an open public product API for commercial monitoring, and that remains the position as of 2026 — there is no self-service, publicly documented product or pricing API; third-party catalogue access exists only through commercial web-data providers that work around the site's own bot-protection layer, not an official ASDA endpoint. Structured extraction from public pages is the practical route. If an official data partnership is available for your use case, pursue that first.
Yes, but they need a separate schema. Clothing and homeware require size, colour and variant fields that grocery does not, and mixing them into a grocery feed will distort category counts and averages. Scope them explicitly, one way or the other.
Yes, with a product matching layer — match on EAN where published, fuzzy-match on brand, title and pack size where not. Own-label lines never match across retailers by identifier and must be matched at category and pack-size level. Crucially, comparisons must be format-controlled: comparing an ASDA Express price against a Tesco superstore price is not like-for-like.
Collecting publicly displayed factual pricing for analysis is a widely practised commercial activity. The risk areas are personal data, database rights, contractual site terms, and conduct that impairs the service. Stay on public pages, avoid personal data, rate-limit conservatively, and take legal advice for your programme.
Actowiz Solutions delivers UK grocery datasets across ASDA and the other major UK retailers, with Rollback and Rewards capture modelled separately, store-format context on every row, own-label tier classification, validated schemas and scheduled delivery to S3, SFTP, BigQuery or API.
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Quick Commerce GCC Dashboard 2026 delivers market, pricing, competitor, assortment, and demand insights for smarter quick-commerce decisions across GCC markets.
Track product availability across Indian pincodes with Zepto/Instamart pincode stock — India for better inventory, assortment, and regional insights.
Actowiz Solutions' 2026 AI training data market report — demand drivers, pricing models, sourcing trends, compliance economics & what data buyers should know.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.