Used vehicle listings
The core dataset, with lifecycle attached.
- Asking price and reduction history
- Mileage, age and owner count
- Condition and service history mentions
- Days on lot with relist linking
- Status transitions with dates
With days-on-lot and reduction history, not just today's asking price.
Two identical cars at the same asking price are not the same asset. One listed yesterday; the other has sat ninety days and reduced twice. That difference is invisible in a snapshot and it is the entire basis of a buying decision.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Automotive data scraping is the automated collection of vehicle listing information from online marketplaces, dealer websites, auction platforms and manufacturer sites: asking price, specification, mileage, condition, trim level, dealer and location.
The category has a familiar shape to real estate, and for the same reason: a vehicle listing is not a static product record, it is an asset with a selling history. That history carries most of the commercial signal.
The same model at two trim levels can differ 20% in value. Listings describe trim inconsistently, with abbreviations, optional packs folded into the title, and market-specific naming. Matching on make and model alone produces price comparisons across genuinely different vehicles.
We match to a normalised trim taxonomy with a confidence score, treat fuel type and transmission as hard constraints, and capture EV-specific fields — battery capacity, quoted range, charging speed — as first-class rather than buried in description text, because in electric vehicles those fields drive price more than trim does.
Transacted prices. Asking prices are public; what a vehicle actually sold for is not, and a listing disappearing does not prove a sale. We report sold_or_withdrawn rather than sold, because those are different outcomes and conflating them corrupts every velocity metric built on top.
Used vehicle pricing is the largest use case. New vehicle configurator pricing is the most technically awkward.
The core dataset, with lifecycle attached.
The attributes that determine value.
Who is holding what, where.
List pricing and configurator structure.
The demand signal layer.
Published payment-led pricing.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 130+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
listing_key |
string | Stable listing identity that survives relisting and source ID changes | Every run |
make / model / trim / trim_confidence |
string / decimal | Normalised specification with a confidence score on trim matching | Weekly |
year / mileage / owners |
int | Registration year, mileage in local units and owner count where published | Weekly |
fuel / transmission |
enum | Fuel type including BEV and PHEV, and transmission, treated as hard match constraints | Weekly |
battery_kwh / range_wltp |
decimal / int | EV battery capacity and quoted range, extracted as fields not description text | Weekly |
asking_price / currency |
decimal / string | Current asking price in local currency | Daily |
price_history / reductions |
array / int | Every observed price change with dates, and a reduction count | Daily |
days_on_lot |
int | Elapsed days from first observation, linked across relistings | Daily |
dealer_id / dealer_type |
string / enum | Stable dealer identity and franchised versus independent classification | Weekly |
lat / lon |
decimal | Dealer or listing geolocation, enabling regional pricing analysis | Weekly |
status |
enum | available, sold_or_withdrawn or relisted — never asserted as sold | Daily |
Status is sold_or_withdrawn rather than sold because a listing disappearing does not prove a transaction. Treating disappearance as a sale inflates every sell-through metric built on the data.
Automotive marketplaces are intensely national. Coverage is built market by market with dealer site depth scoped separately.
Vehicle identification numbers are collected only where a source publishes them openly. We do not query registration or DVLA-type lookup services to enrich listings, since those carry their own terms and often personal data linkage. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United Kingdom & Germany | The deepest used vehicle marketplaces with rich published specification, which makes trim matching and lifecycle tracking unusually reliable. |
| United States | Enormous dealer base and strong online retail penetration, with heavy demand for residual value evidence from lenders. |
| France, Spain & Netherlands | Well-developed marketplaces plus fast EV transition, driving demand for battery and range field coverage. |
| India & United Arab Emirates | Rapidly growing organised used car retail with high listing churn, where relist detection matters most. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Dealers and marketplaces dominate, with lenders, remarketers and OEM pricing teams close behind.
Stock pricing decisions need to know how comparable vehicles are priced locally and how long they are actually sitting.
Comparable vehicle pricing by trim and mileage band within a radius, with days on lot and reduction history on every comparable.
Stock turn days
Sourcing decisions need to know which segments are moving and where vehicles are mispriced enough to buy.
Segment sell-through velocity plus listings with high days on lot and reduction history, filtered to your buying criteria.
Gross per unit
Your pricing guidance and search ranking need comprehensive competitor listing coverage with normalised specification.
Normalised listing data across competing marketplaces with trim matching and lifecycle fields, delivered to your systems.
Listing quality score
Residual value models need current asking price evidence by trim and age, not lagged guide book values.
Longitudinal asking price panels by make, model, trim, age and mileage band with reduction behaviour included.
RV forecast accuracy
Defleet timing and channel choice depend on current retail pricing and how quickly segments are selling.
Retail asking price and sell-through velocity by segment and region, so defleet timing is evidence-based.
Disposal proceeds
Automotive retail theses need observable inventory, pricing and velocity data rather than quarterly commentary.
Longitudinal inventory, pricing and velocity panels by retailer, segment and market for direct modelling.
Signal lead time
Four patterns, with the outcome each is judged on.
For each vehicle in stock, comparables are assembled by trim, age and mileage band within a configurable radius, with days on lot and reduction history on every comparable so pricing reflects what is actually selling rather than what is merely listed.
Outcome: Pricing set against genuinely comparable local stock instead of guide book values alone.
Listings with high days on lot and repeated reductions are surfaced against your buying criteria, since that population carries the most negotiation leverage.
Outcome: Buying pipelines built from evidenced dealer pressure rather than from asking prices alone.
Asking prices are tracked by make, model, trim, age and mileage band over time, including reduction behaviour, producing current market evidence ahead of published guide revisions.
Outcome: Residual assumptions tested against observed retail asking prices rather than lagged guides.
Listing volume, sell-through velocity and inventory ageing are tracked by segment, including BEV versus ICE mix shift with battery and range fields available for EV-specific analysis.
Outcome: Demand shifts visible by segment months before registration data reflects them.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Comparable pricing analysis used marketplace days-on-lot figures, which reset whenever a dealer relisted, so genuinely stale competitor stock appeared newly listed.
Relist detection using specification, mileage, dealer and image fingerprinting, with days_on_lot continued from original first-seen and reduction history joined.
Comparable analysis began reflecting real market age, changing pricing decisions on slow-moving segments.
Residual modelling relied on published guide values, which trailed observable retail movement during a period of rapid EV price change.
Longitudinal asking price panels by make, model, trim, age and mileage band, including reduction behaviour and EV battery and range fields.
Residual assumptions were tested against current retail evidence ahead of guide revisions.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Trim normalisation and lifecycle identity across relistings are the parts in-house builds consistently lack.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Days on lot is the most useful field in automotive data and the easiest to get wrong. The reason is behavioural: dealers relist ageing stock to reset the visible counter and regain search prominence.
Listings are matched across relists on specification, mileage, dealer and image fingerprinting. Mileage is a strong signal here: a relisted vehicle usually carries the same or near-identical mileage, which distinguishes it from a genuinely different unit of the same trim.
Image fingerprinting matters because dealers often reuse the same photography. Where a relist is detected, days_on_lot continues from the original first-seen date, and the reduction history is joined rather than restarted.
Confidence is scored, and where linking is uncertain we flag it rather than merging two vehicles that might be different. In a category where mileage and trim can genuinely coincide, a wrong merge is worse than an acknowledged uncertainty. The same reasoning applies in our real estate service, where relisting behaviour is near-identical.
Every automotive data buyer eventually asks for transacted prices. It is the right thing to want and it is not publicly available.
A listing disappearing can mean sold, withdrawn unsold, moved to auction, or moved to another platform. We report sold_or_withdrawn and refuse to assert a sale, because inflating sell-through is the single easiest way to make this dataset misleading while looking more useful.
If your model genuinely requires transacted prices, the routes are auction data licences, dealer management system partnerships or registration-linked datasets — all licensed, none obtainable by scraping. We will point you there rather than sell a proxy dressed up as the real thing.
Markets, sources and whether dealer site coverage is needed are scoped first, since dealer sites are more work than marketplaces.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file. Geocoded listings support radius-based comparable analysis directly in PostGIS or BigQuery GIS.
We collect publicly accessible vehicle listing pages from marketplaces, dealer sites and manufacturer configurators. We do not query registration lookup services, access dealer management systems, or collect private seller personal contact details. VINs are collected only where a source publishes them openly.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What dealer, marketplace and lending teams ask during evaluation.
No, and no scraping vendor can. Transacted prices are not publicly published, and a listing disappearing does not prove a sale — it can equally mean withdrawn unsold, moved to auction, or moved to another platform.
We report sold_or_withdrawn and refuse to assert a sale. If your model requires transacted data, the routes are auction data licences, dealer management system partnerships or registration-linked datasets. All are licensed, and we will point you there rather than sell you a proxy presented as the real thing.
Relists are detected and linked, so days_on_lot continues from the original first-seen date rather than restarting. Matching uses specification, mileage, dealer and image fingerprinting — mileage is especially strong, since a relisted vehicle carries near-identical mileage.
This is the field most clients build their analysis on, and without relist linking it is actively misleading: ageing stock appears fresh, which removes exactly the signal you bought the data for.
Around 96% to a normalised trim taxonomy, with a confidence score on every listing. Fuel type and transmission are treated as hard constraints rather than soft signals.
Trim matters enormously here — the same model across two trim levels can differ 20% in value — and listings describe it inconsistently with abbreviations and optional packs folded into titles. Below your chosen confidence threshold, records arrive flagged rather than force-matched.
Yes, as first-class fields rather than description text: battery capacity in kWh, quoted WLTP or EPA range, charging speed where published, and battery health or state-of-health claims where a listing includes them.
In electric vehicles these fields drive price more than trim does, so leaving them in free text makes the dataset far less useful. Where a listing does not publish them, the field is null rather than estimated from the model name.
Yes, and dealer sites are often more valuable because stock appears there before or instead of marketplace syndication, and finance offers are usually published there.
They are also considerably more work: thousands of individual sites on many different platforms. We scope dealer coverage separately from marketplace coverage and prioritise the dealer groups that matter to you, rather than attempting exhaustive coverage that would inflate cost for marginal benefit.
Only where a source publishes them openly, which some markets and platforms do. We do not query registration or lookup services to enrich listings, because those carry their own terms and frequently link to keeper or owner personal data.
Where VINs are absent, our listing identity is built from specification, mileage, dealer and image fingerprinting, which is sufficient for relist linking and comparable matching in practice.
Yes, including manufacturer list prices by trim, option and pack pricing, and published finance representative examples. Configurator collection is technically awkward because options are interdependent and pricing changes as selections are made.
We collect the published price structure rather than attempting to enumerate every possible configuration, which would produce combinatorial volume for little analytical gain. Where a manufacturer publishes campaign discounts, those are captured as separate fields with dates.
We collect the vehicle and pricing data, and we do not collect private seller names, phone numbers or contact details. Those are personal data belonging to individuals rather than business information.
Private listings are useful for market analysis because they price differently from dealer stock, so excluding them entirely would distort segment pricing. Dealer versus private classification is delivered as a field so you can separate them in analysis.
We quote individually. The drivers are market count, whether dealer site coverage is needed alongside marketplaces, listing volume, and refresh frequency — plus whether lifecycle tracking with history reconstruction is required, since that needs continuous coverage of the same listing population.
A defined market and segment across major marketplaces at daily refresh sits at the lighter end. Multi-market coverage including thousands of dealer sites with sub-daily new-listing alerting sits higher. One scoping call, a free pilot on your own market within 48 hours, then a fixed monthly quote. Request a quote.
Send us a market, segment or stock list. We return listings with trim matching, days on lot and reduction history within 48 hours.
Free pilot, no card, no obligation. We'll be clear that asking prices are not transacted prices.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
The EU AI Act impact on web scraping & AI training data GPAI transparency, copyright reservations, prohibited practices & a compliance checklist from Actowiz.
How a B2B supplier replaced manual tender-portal checking with an automated, filtered feed of relevant government tenders from GeM and CPP/eProcure never missing a bid deadline again.
Actowiz Solutions tracks post–World Cup 2026 travel pricing — hotel ADR & airfare normalization across host cities, event-premium decay data & lessons for travel teams.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.