Aldi UK is a completely different data problem from the big four. The core grocery range is small and tightly curated rather than tens of thousands of SKUs, so catalogue scale is not the challenge. Two things are. First, Specialbuys are ephemeral — time-limited, non-grocery lines that drop on a fixed schedule, sell out, and disappear. If you are not capturing during the window, that data is gone permanently and cannot be backfilled. Second, almost everything Aldi sells is own label, which means EAN-based matching against Tesco or Sainsbury's largely does not work, and cross-retailer comparison requires an attribute-based matching layer instead. This guide covers both problems, the field schema, a sample dataset, and the build-versus-buy maths.
Aldi is the benchmark. That is the single most important commercial fact about this dataset.
Sainsbury's runs Aldi Price Match. Tesco has run price matching against the discounters. ASDA positions Just Essentials against them. When the big four make competitive pricing decisions, Aldi is frequently the reference point they are measuring against [VERIFY — cite current price matching schemes and their published scope]. That makes Aldi pricing disproportionately valuable relative to its market share [VERIFY — cite Kantar Worldpanel grocery market share, latest reading], because it is the input to everyone else's pricing strategy.
But Aldi's operating model makes the data harder to obtain and harder to use than any of the big four, for three structural reasons.
The range is deliberately small. Aldi's business model is built on range discipline — a tightly curated assortment measured in the low thousands of SKUs rather than the tens of thousands a big-four supermarket carries [VERIFY — confirm current approximate Aldi UK core range size]. Fewer SKUs, less shelf complexity, lower cost to serve. For a data team, this means catalogue scale is not the problem. Depth of coverage on each individual line is.
Online grocery presence is narrower than the big four. Aldi UK's e-commerce operation has historically focused on non-food and Specialbuys, with grocery available through click and collect rather than the full national delivery model Tesco or Sainsbury's run [VERIFY — confirm current Aldi UK online grocery availability and coverage before publishing]. This is the honest constraint in this article: the volume and structure of pricing data published online for Aldi grocery differs from what the big four publish, and any credible data partner should tell you that upfront rather than implying identical coverage.
Almost everything is own label. The overwhelming majority of Aldi's range is exclusive own-brand product [VERIFY — confirm current own-label proportion]. That has a consequence most teams do not anticipate until their matching layer fails, covered in detail below.
Specialbuys is Aldi's rotating non-grocery range — the middle-aisle proposition. New lines drop on a fixed weekly schedule, typically Thursdays and Sundays [VERIFY — confirm current Specialbuys drop schedule], in limited quantity, while stocks last. Popular lines sell out within hours.
For a data pipeline, this changes everything about scheduling.
With a big-four grocery catalogue, a product that appears today will almost certainly still be there tomorrow. Miss a daily capture and you lose one observation of a continuing series. With Specialbuys, a product can appear, sell out and disappear between two daily captures. If your schedule is daily at 6am and a line drops at 8am and sells out by 2pm, that product never existed as far as your dataset is concerned.
You cannot backfill it. There is no archive to go back to. The data is simply gone.
This has concrete implications:
Cross-retailer price comparison normally works like this: match products on EAN, compare prices on identical items, report the gap. Clean and defensible.
That approach largely fails on Aldi, because an Aldi own-brand product has no equivalent EAN at Tesco. There is no identical item. A shopper comparing a basket is comparing an Aldi own-label baked bean against a Heinz baked bean and a Tesco own-label baked bean — three different products serving the same need at three different quality and price positions.
So Aldi comparison requires attribute-based matching, which is materially harder and needs to be honest about its own uncertainty:
This is where a lot of published "Aldi is X% cheaper" analysis falls apart under scrutiny. If the matching methodology is not stated and the confidence not exposed, the number is not defensible. Any serious Aldi comparison dataset ships its matching logic alongside its prices.
| Field | Description | Example |
|---|---|---|
| product_id | Aldi internal product identifier | 4088600123456 |
| product_url | Canonical product page URL | https://www.aldi.co.uk/... |
| product_name | Full product title as displayed | Example Brand British Semi Skimmed Milk 2.27L |
| brand | Aldi exclusive brand name | Cowbelle |
| is_own_label | Boolean — Aldi exclusive brand | true |
| range_type | Core grocery or Specialbuy | core / specialbuy |
| own_label_tier | Tier where applicable | Everyday Essentials / core / Specially Selected |
| pack_size | Size, weight or volume as listed | 2.27L |
| gtin_ean | Barcode identifier, where published | 4088600123456 |
| category_path | Full breadcrumb hierarchy | Fresh Food > Dairy & Eggs > Milk |
| image_urls | Array of product image URLs | ["https://www.aldi.co.uk/..."] |
Aldi's own-label architecture runs from Everyday Essentials at the value end, through the core exclusive brands, to Specially Selected at the premium end. Capturing the tier is essential — it is the only way to attempt credible equivalence matching against big-four own-label tiers.
Note that brand for Aldi is usually an exclusive brand name rather than a national brand. Carry both brand and is_own_label, because an analyst filtering for "own label" cannot do it on brand name alone when the brand names look like real brands.
| Field | Description | Example |
|---|---|---|
| price | Current price | 1.45 |
| currency | ISO currency code | GBP |
| pricing_basis | Per item or per weight | per_item / per_kg |
| unit_price | Normalised price per standard unit | 0.64 |
| unit_of_measure | Basis for the unit price | per litre |
| was_price | Previous price where a reduction is shown | 1.69 |
| promo_type | Nature of the offer where present | price_drop / super_6 / multibuy |
| promo_text | Raw offer text exactly as displayed | Super 6. Was £1.69 |
| is_super_six | Boolean — part of the rotating fresh promotion | false |
Aldi's promotional structure is deliberately simpler than the big four's. There is no loyalty price layer, because there is no loyalty card mechanic of the Clubcard or Nectar type. That simplicity is a genuine advantage in the schema — one price means one price.
Super 6 is the recurring rotating produce and meat promotion, and it is worth flagging separately because it is the closest thing Aldi has to a regular promotional cycle, and tracking which lines rotate through it over time is a useful assortment signal.
| Field | Description | Example |
|---|---|---|
| is_specialbuy | Boolean — rotating limited range | true |
| on_sale_date | Published date the line goes on sale | 2026-03-12 |
| first_seen_at | First capture where this product appeared | 2026-03-12T08:04:00Z |
| last_seen_at | Most recent capture where it was present | 2026-03-12T16:22:00Z |
| sold_out_at | Derived — first capture showing sold out | 2026-03-12T17:00:00Z |
| availability_duration_hours | Derived — observed time on sale | 9.0 |
| stock_status | Current state | available / low_stock / sold_out / window_ended |
| specialbuy_theme | Themed drop where applicable | Garden & Outdoor |
first_seen_at, sold_out_at and availability_duration_hours are derived from observation, not published. They cannot be reconstructed later. This is the clearest example in the entire UK grocery cluster of data that either gets collected at the moment it exists or does not exist at all.
| Field | Description | Example |
|---|---|---|
| availability_status | Stock state at time of capture | in_stock / out_of_stock / sold_out |
| store_context | Location context for this capture | national / store identifier |
| fulfilment_type | Where applicable | click_collect / delivery / in_store_only |
| captured_at | UTC timestamp of the capture | 2026-03-12T08:04:00Z |
An illustrative record showing the output schema for a Specialbuy. Values are synthetic, shown to demonstrate field shape and types — they do not represent live Aldi pricing. Request a live sample for real current data.
{
"product_id": "4088600123456",
"product_url": "https://www.aldi.co.uk/example-product",
"product_name": "Example Cordless Pressure Washer",
"brand": "Workzone",
"is_own_label": true,
"range_type": "specialbuy",
"own_label_tier": null,
"pack_size": null,
"gtin_ean": "4088600123456",
"category_path": "Specialbuys > Garden & Outdoor > Cleaning",
"price": 79.99,
"currency": "GBP",
"pricing_basis": "per_item",
"unit_price": null,
"was_price": null,
"promo_type": null,
"is_super_six": false,
"is_specialbuy": true,
"on_sale_date": "2026-03-12",
"first_seen_at": "2026-03-12T08:04:00Z",
"last_seen_at": "2026-03-12T16:22:00Z",
"sold_out_at": "2026-03-12T17:00:00Z",
"availability_duration_hours": 9.0,
"stock_status": "sold_out",
"specialbuy_theme": "Garden & Outdoor",
"availability_status": "sold_out",
"captured_at": "2026-03-12T17:00:00Z"
}
A core grocery line, flattened to CSV alongside Specialbuys:
| product_id | product_name | brand | range | tier | price | unit_price | super_6 | status | duration_hrs |
|---|---|---|---|---|---|---|---|---|---|
| 4088600123456 | Cordless Pressure Washer | Workzone | specialbuy | — | 79.99 | — | false | sold_out | 9.0 |
| 4088600123457 | Semi Skimmed Milk 2.27L | Cowbelle | core | core | 1.45 | 0.64/L | false | in_stock | — |
| 4088600123458 | Value Baked Beans 410g | — | core | Everyday Essentials | 0.29 | 0.07/100g | false | in_stock | — |
| 4088600123459 | Aged Cheddar 350g | Specially Selected | core | Specially Selected | 3.29 | 0.94/100g | false | in_stock | — |
| 4088600123460 | Loose Broccoli | — | core | core | 0.69 | — | true | in_stock | — |
Two things this table shows.
The Specialbuy row has a duration_hrs value and the core rows do not, because duration only means something for a time-limited line. Nine hours from first appearance to sold out — that single number tells a buyer more about demand than the price does, and it exists only because the capture schedule was dense enough to observe both ends of it.
Rows three and four show the tier spread: an Everyday Essentials line at £0.29 and a Specially Selected line at £3.29. Any credible comparison against a big-four retailer has to match tier to tier. Comparing an Everyday Essentials bean against a Tesco Finest bean produces a headline number that is technically accurate and completely meaningless.
Read https://www.aldi.co.uk/robots.txt and honour it. Confine collection to publicly accessible pages — no logged-in areas, no account data, no personal data. If a path is disallowed, it is out of scope.
This matters more than usual on Specialbuys drop days. Drop mornings are Aldi's highest-traffic moments, and it is exactly the wrong time to point aggressive collection at their infrastructure. Conservative rates on drop days are both an ethical and a practical requirement — heavy load during a launch window is the fastest way to end a data programme permanently.
For most retailers, the scraper is the interesting part and the scheduler is plumbing. For Aldi Specialbuys, it is the reverse.
A workable pattern:
# Illustrative scheduling logic — the interesting part of an Aldi pipeline
SPECIALBUY_DROP_DAYS = {3, 6} # Thursday, Sunday (0 = Monday) — VERIFY current schedule
def capture_interval_minutes(now) -> int:
"""Dense sampling during drop windows, sparse otherwise."""
if now.weekday() not in SPECIALBUY_DROP_DAYS:
return 720 # twice daily on non-drop days
hours_since_open = now.hour - 8 # assume drop at store-open
if hours_since_open < 0:
return 60 # hourly pre-drop, to catch early listings
if hours_since_open <= 4:
return 15 # dense through the sell-through curve
if hours_since_open <= 10:
return 60
return 240
The point is not the exact numbers — tune them to observed behaviour. The point is that a flat daily schedule cannot produce sell-out timing, and sell-out timing is the most valuable field in an Aldi Specialbuys dataset.
sold_out_at and availability_duration_hours are derived from a state transition between captures. That makes them vulnerable to gaps in a way that captured fields are not.
If a capture fails at 2pm and the next succeeds at 5pm, and the product sold out at 2:30pm, your derived duration is wrong by up to three hours — and nothing in the data indicates that. Derived fields therefore need a companion:
Extract from structured data where it exists rather than CSS selectors — many retail pages publish Product schema in JSON-LD, which changes far less often than front-end class names. Identify your collector honestly in the user agent and rate-limit conservatively.
Validation specific to Aldi:
Formats. CSV and Excel for commercial teams. JSON or JSONL for pipelines. Parquet where volume and query cost matter — though Aldi volumes are modest compared with big-four feeds.
Destinations. S3, Google Cloud Storage or Azure Blob; SFTP; direct load into BigQuery, Snowflake or Redshift; or a REST endpoint.
Delivery shape. Aldi splits naturally into two streams, and they should be delivered separately.
Alerting. For competitors and category buyers tracking Aldi, the alert is the product. A new Specialbuy drop in a monitored category is worth knowing about within the hour, not in Friday's file.
The big four grocers are the largest buyers of Aldi price data, because Aldi is the benchmark their price-matching schemes reference. Anyone operating an Aldi Price Match programme needs continuous, auditable Aldi pricing to administer it credibly.
CPG and FMCG brands track Aldi own-label lines as competitive substitutes. A brand losing share to an Aldi exclusive needs to see the price gap and the tier positioning, not just a headline number.
General merchandise and homeware brands track Specialbuys specifically. An Aldi Specialbuy at £79.99 in a category where a brand retails at £150 is a direct competitive event with a known start date and a known end. The sell-out speed tells them how much demand there was at that price.
Price comparison platforms and consumer publications compare baskets. This audience needs the matching methodology and confidence scores more than anyone, because their output gets scrutinised publicly.
Analysts and researchers use Aldi as the value anchor in UK food inflation work, and use multi-year Specialbuys archives to study discounter assortment strategy — a dataset that cannot be built retrospectively.
Public data only. Collect what any visitor can see without authenticating. No logged-in pages, no account areas, no personal data.
Load discipline during drop windows. This is the Aldi-specific ethical point. Dense sampling during a Specialbuys launch is legitimate data collection; hammering the site during the busiest hour of its week is not. Keep per-capture volumes small, space requests, and back off immediately on any sign of degraded response.
No personal data. Prices are not personal data. Reviews may contain reviewer names or identifiable content — if you collect reviews, UK GDPR applies and you need a lawful basis, retention policy and minimisation. For price monitoring, collect counts and averages only.
Database rights. The UK retains a sui generis database right, separate from copyright, protecting substantial investment in obtaining, verifying or presenting database contents. Extracting a substantial part can infringe it. The defensible position is factual price monitoring for analysis and comparison.
Comparative claims. If your output supports public "cheaper than" claims, the matching methodology carries legal weight as well as analytical weight. Misleading comparative advertising is regulated in the UK. Ship the methodology and the confidence scores.
Terms of service. Site terms are contractual and enforceability against non-account-holders varies. Treat them as a real consideration.
Not legal advice. Take advice from a qualified UK solicitor for your specific programme.
Aldi inverts the usual build-versus-buy logic, and it is worth being clear about why.
The core range is the easiest in-house build in UK grocery. A few thousand SKUs, no loyalty price layer, simple promotional structure. A competent engineer can stand up a daily core-range capture in a week and maintain it with modest effort.
Specialbuys is the hardest. Not because the parsing is difficult — it is not — but because the operational requirement is unforgiving. Dense sampling on a fixed schedule, twice a week, every week, with immediate alerting on capture failure, because a missed window is permanent data loss. That is not a project. It is an on-call commitment, and in-house teams almost always underestimate what it costs to sustain over two years.
The honest framing: if you only need core range pricing, build it. If you need Specialbuys lifecycle data, buy it — or accept that you will have gaps, and be upfront with your stakeholders about which weeks are missing.
No. Specialbuys are time-limited and sell out, and there is no public archive to recover them from. If a product was not captured during its window, it does not exist in the dataset and cannot be added later. This is the strongest argument for starting collection before you think you need it.
Core range twice daily is sufficient — the range is stable and promotions are simple. Specialbuys need dense sampling aligned to drop days, with intervals as short as 15 minutes through the first hours to observe sell-through, tapering afterwards. A flat daily schedule cannot produce sell-out timing.
Because the overwhelming majority of Aldi's range is exclusive own-brand product with no equivalent item at other retailers. Comparison requires attribute-based matching on category, pack size, normalised unit price and own-label tier equivalence, with a confidence score exposed on every matched pair.
Aldi UK's e-commerce operation has historically been narrower than the big four, focused on non-food and Specialbuys with grocery available through click and collect rather than full national delivery [VERIFY — confirm current position before publishing]. Any data partner should be upfront about coverage differences rather than implying parity with Tesco or Sainsbury's.
Yes, derived from the transition between an "available" capture and a "sold out" capture — but only with dense enough sampling, and the observation gap should always be shipped alongside the duration as an uncertainty band. A duration without its uncertainty is not a defensible number.
Collecting publicly displayed factual pricing for analysis is a widely practised commercial activity. The risk areas are personal data, database rights, contractual site terms, and conduct that impairs the service — the last of which deserves particular care during Specialbuys drop windows. Take legal advice for your specific programme.
Actowiz Solutions delivers UK grocery datasets across Aldi and the other major UK retailers, with Specialbuys lifecycle capture on drop-aligned schedules, own-label tier classification, attribute-based cross-retailer matching with exposed confidence scores, validated schemas and scheduled delivery to S3, SFTP, BigQuery or API.
You can also reach us for all your mobile app scraping, data collection, web scraping , and instant data scraper service requirements!
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Extract Aldi UK product, price and Specialbuys data at scale. Capture windows for time-limited drops, own-label matching, field schema and UK compliance.
Multi-Channel Marketplace Inventory Scraping API helps brands monitor product stock, availability, and inventory changes across Amazon, Flipkart, and Myntra.
Sephora & Trendyol Arabic Market Data Report 2026 delivers UAE e-commerce intelligence on products, pricing, trends, and customer demand.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.