Vivino Wine Matching Dataset, Twice a Week
This is a matching feed rather than a catalogue dump: eight columns that take your wine SKUs and hand back the wine's page on Vivino, the largest wine community in the world.
Twice a week suits a stable list: two runs in seven days keep a catalogue matched at the lowest cost, with the same eight columns and the same identifiers as the faster feeds.
The file already exists, because this pipeline runs twice a week whether you buy it or not. Most vendors start collecting after you order — which is why they quote a lead time. Here the most recent file is in your account within minutes of payment, with an API key issued at the same time.
What you get, in plain terms
Four things. The same matching feed, at the cadence a settled list needs.
Keyed on your own SKU
Sku, ProductId and VariantId are your catalogue's identifiers, carried through unchanged, so the file joins to your product master with no name matching at all.
The Vivino page as the output
Vivino Weblink is the matched wine's page on Vivino — the link behind the rating, the tasting notes and the community reviews.
Bottle volume on the row
BottleVolume gives 750, 500, 375 or 1500 in millilitres, so half bottles and magnums are matched as separate wines rather than collapsed into the standard size.
However you want it
Direct download, REST API, Amazon S3, Google Cloud, Snowflake or SFTP. The API key comes with the dataset.
Fields included in this dataset
All 8 columns, in the exact order they appear in the file — taken straight from the sample, not from a brochure. The free sample ships with a data dictionary giving an example value for each one.
Sample rows from the real file
Real rows from the sample file, not an illustration. The supplied sample has 50 SKU rows with all 8 columns.
Coverage
Captured twice a week. One row is one wine in your catalogue, with the identifiers you supplied carried through unchanged.
| Metric | In the supplied sample | What it tells you |
|---|---|---|
| Rows | 50, all distinct SKUs | One row per catalogue line |
| Columns | 8 | A matching feed, not a catalogue dump |
| Bottle volumes | 750 ml mostly · 375 · 500 · 1500 | Half bottles and magnums kept apart |
| Non-vintage wines | Around half the sample | “NV” written into the name |
Vivino Weblink |
“Not Found” on all 50 rows | See the note below |
Brand |
Empty on all 50 rows | Producer sits inside the name |
Matching a larger list, or only part of it? Send the SKUs and we will price the run.
Historical data
The last 30 days come with the dataset. Beyond that we hold Vivino match records from February 2025 onwards.
Ask about historical dataNeed more data points?
We can extend this dataset beyond the standard 8 columns — the Vivino rating, review count and region pulled onto the row are the most requested additions, along with a confidence score on each match.
Request custom fields“Our range barely changes between deliveries. Two runs a week keeps everything matched and costs a third less.”
What people use this dataset for
Wine retailers and e-commerce
Put a recognised rating next to every bottle on your own site, matched line by line.
Importers and distributors
Check how the wines you carry are represented on the platform your customers check first.
Wine apps and marketplaces
Enrich a supplier feed with a canonical wine reference instead of matching names yourself.
Data teams
Use the matched pairs as training data for wine entity resolution, where names, vintages and volumes all vary.
About Vivino matching data
Vivino is the world's largest wine community, with ratings and reviews on millions of labels. Shoppers check it before buying, which makes the link between your catalogue line and its Vivino page commercially useful — and surprisingly hard to produce, because wine names, vintages and bottle sizes rarely agree between systems.
Your identifiers survive the round trip
Sku, ProductId and VariantId come back exactly as supplied, so the file joins straight to your product master. BottleVolume keeps 375 ml, 500 ml and 1.5 l bottlings separate from the standard 750.
Why twice a week
If your range turns over slowly, an hourly or daily run mostly re-matches wines that have not changed. Two captures a week keep the catalogue current and leave the budget for the wines that are genuinely new.
The supplied sample is a coverage warning
Every one of the 50 sampled rows came back with Vivino Weblink set to “Not Found”, and Brand was empty throughout. These are small-production and private-label wines — exactly the part of a catalogue that a community database is least likely to hold — so the sample shows the schema and the identifiers working, not the match rate you would see on mainstream labels. Ask us for a sample built on your own SKU list before you buy: match rate is the only number that matters in a feed like this, and it depends entirely on which wines you sell.
Is collecting this data legal?
We collect only what is publicly visible on a wine listing — producer, region, grape, vintage, price, availability and the retailer's own description. No account is created, no age gate is circumvented and no personal data is touched, so the file holds no consumer records. These are age-restricted products: the data describes what is listed publicly, and what you may do with it depends on your own market's rules. Producer names, tasting notes and imagery stay the property of their owners; the licence covers the compiled dataset.
