Shade & variant availability
The layer that makes the data meaningful.
- Per-shade availability with full curve
- Shades available against total ranged
- Broken shade range detection
- Deep-shade stockout flags
- Shade introduction and discontinuation
At shade level, because a foundation with two shades left is not in stock.
Beauty has the same structural problem as fashion and almost nobody treats it that way. A foundation ranges forty shades. Product-level data says it is in stock. Thirty-eight shades are gone, which is the opposite of what the data implies.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Beauty and personal care data scraping is the automated collection of cosmetics, skincare, haircare and personal care retail data: pricing and unit pricing, per-shade availability, ingredient lists, published claims, sampling mechanics, assortment and category placement.
The category has the same structural problem as fashion, and almost no dataset treats it that way: shade is the purchasable unit, not the product.
Every record is a product-retailer combination carrying a full shade curve: per-shade availability, count available against total ranged, plus derived flags. shade_range_broken indicates gaps across the range. deep_shades_oos flags whether the deeper end of the range is unavailable, which is the field most brand and category teams actually came for.
Shade naming is chaotic — numeric codes, descriptive names, or both, with no cross-brand standard. We parse and retain exactly what is published rather than forcing a standardised scale that does not exist.
Verify claims. Published claims like "vegan", "non-comedogenic", "dermatologist tested" or an SPF figure are captured verbatim with claims_verified set to false. We are not a testing house, and presenting a captured claim as validated would be the most damaging thing we could do in this category.
Shade availability is the field clients most often did not know they needed. Ingredient capture is the fastest growing.
The layer that makes the data meaningful.
Comparable across pack sizes.
INCI as structured data.
Captured verbatim, never verified.
How beauty actually promotes.
Range and newness.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 110+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
product_key / brand |
string | Cross-retailer product identity and normalised brand | Every run |
shade_curve |
array | Per-shade availability, the field that makes shade-level demand visible | Daily |
shades_available / shades_total |
int | Purchasable shades against total ranged at that retailer | Daily |
shade_range_broken / deep_shades_oos |
boolean | Derived flags for range gaps and deep-shade unavailability | Daily |
price / unit_price_computed / unit_basis |
decimal / string | Price and unit price computed by us on a consistent basis | Daily |
inci_captured / inci_list / inci_count |
boolean / array / int | Whether the ingredient list was published, and its content | Weekly |
claims_published / claims_verified |
array / boolean | Claims verbatim, with verified always false | Weekly |
gwp_offer / sample_included |
string / boolean | Gift-with-purchase mechanics and sample inclusion as displayed | Daily |
is_exclusive / is_limited_edition |
boolean | Retailer exclusivity and limited edition status where indicated | Weekly |
first_seen / discontinued_at |
date | Launch observation and discontinuation detection | Daily |
shade_added_at |
date | When a new shade entered the range, for range expansion tracking | Daily |
claims_verified is a constant reading false. It exists so nobody downstream mistakes a captured claim for a validated one. We record what a brand published; whether it is substantiated is a testing question, not an extraction one.
Beauty retail is a mix of specialists, department stores, pharmacy and brand D2C. Coverage is built to your competitive set.
Shade-level availability is published by most beauty specialists and by fewer general retailers. We confirm per retailer which expose shade availability publicly before build rather than substituting product-level stock. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United Kingdom & France | The deepest beauty specialist retail with widespread public shade-level availability, which makes shade curve analysis most complete. |
| United States | Largest prestige beauty market with the most scrutinised deep-shade ranging, and heavy launch cadence. |
| India | Fast-growing beauty ecommerce on Nykaa, Purplle and Tira with aggressive sampling and bundle mechanics. |
| United Arab Emirates & Southeast Asia | High prestige beauty penetration with distinct shade range requirements and strong travel retail overlap. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Brand and retailer category teams dominate, with formulation and regulatory close behind.
Range decisions need to know shade-level sell-through at competitors, which product-level stock data cannot show.
Shade curves with broken-range and deep-shade flags across your competitive set, plus assortment breadth by brand.
Range productivity
Retail partners run out of your core shades and nobody notices until sales dip.
Per-retailer shade availability on your products with deep-shade stockout flags and replenishment detection.
Availability on core shades
Competitor formulation shifts are visible in INCI lists and nobody is reading them systematically.
INCI lists captured with change detection, so reformulation and ingredient trend adoption become measurable.
Time to competitive insight
Claim substantiation review needs claims as worded, with dates, across the market.
Published claims captured verbatim with wording change detection and first-seen dates.
Claim review coverage
Gift-with-purchase and sampling drive beauty promotion and are not tracked as mechanics.
GWP thresholds, sample inclusion and bundle composition captured as displayed, with validity windows.
Promotional ROI
Beauty theses need observable launch performance and range signals ahead of reported results.
Longitudinal shade availability, launch cadence and assortment panels by brand and retailer.
Signal lead time
Four patterns, with the outcome each is judged on.
Per-shade availability is tracked daily, so products losing core shades while remaining listed are identified as strong sellers, and shades returning to stock indicate replenishment.
Outcome: Launch and range performance read from shade movement rather than from a product-level stock flag.
The deeper end of every shade range is tracked separately with its own stockout flag, per retailer.
Outcome: Deep-shade ranging and replenishment measured across the market rather than asserted.
INCI lists are captured with change detection, revealing reformulations and the adoption or removal of specific ingredients across competitor ranges.
Outcome: Competitor formulation shifts detected from published ingredient lists rather than from press coverage.
Published claims are captured verbatim with first-seen dates and wording change detection, including quiet withdrawals.
Outcome: A dated claim inventory across the market, which is what a substantiation or regulatory review starts from.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Availability monitoring ran at product level, so a foundation with two of forty shades remaining reported as in stock.
Shade-level collection with per-shade curves, deep-shade stockout flags and replenishment detection across retail partners.
Core and deep shade stockouts became visible per retailer within a collection cycle.
The category team had no systematic view of competitor ingredient changes, so reformulations surfaced weeks late.
INCI list capture with formulation change detection across the competitive range, refreshed weekly.
Reformulations and ingredient adoption became measurable from published lists rather than from announcements.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Shade parsing across inconsistent naming conventions is the part in-house builds consistently underestimate.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
When beauty clients see shade-level data for the first time, the question they ask within minutes is always the same: what is happening at the deep end of the range.
We flag deep_shades_oos separately from the general broken-range flag, and we distinguish not ranged from ranged but unavailable — the same distinction that matters in quick commerce between listed and in stock. A shade never ranged at a retailer is a category conversation; a ranged shade persistently out is a supply conversation.
Defining the deep end requires judgement, since shade naming has no cross-brand standard. We derive it per brand and product line from the published range ordering rather than applying a fixed global rule, and we state the method in the scope document so the flag is inspectable rather than a black box.
Beauty is dense with claims, and the temptation to sell "claim verification" is obvious. We do not, and the boundary is worth stating clearly because it protects you as much as us.
Assess whether a claim is substantiated, whether an SPF figure is accurate, or whether an ingredient list is complete. Those require laboratory testing and regulatory expertise. claims_verified is a constant false so that no downstream system can mistake capture for validation.
Nor will we give formulation or dermatological advice on the data. We provide the published ingredient list; interpreting it for safety or efficacy is your regulatory and R&D function's work, and a data vendor offering that opinion would be creating liability for you.
For claim monitoring across sustainability specifically, our ESG service applies the same capture-not-verify principle.
Retailers, categories and whether shade-level collection is required are scoped first, since shade curves multiply record volume substantially.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect publicly accessible product, category and search pages. Published claims and ingredient lists are captured as published and never verified, with claims_verified constant false. We provide no formulation, safety or dermatological assessment. Reviewer names and profiles are not part of the deliverable.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What category, brand and regulatory teams ask during evaluation.
Because shade is the purchasable unit. A forty-shade foundation with two shades left reads as in stock at product level, when it has sold out of thirty-eight.
Product-level data also cannot answer the question this category most cares about — whether deep shades are ranged and replenished. That is invisible without a shade curve, and it is usually the first thing clients ask once they see the data.
No, categorically. We capture claims verbatim with claims_verified set permanently to false. Verification requires laboratory testing and regulatory expertise, and presenting a captured claim as validated would be the most damaging thing we could do in this category.
What we do provide is a dated claim inventory with wording change detection — including quietly withdrawn claims, which is often the most informative event. That is what a substantiation review starts from.
Yes, where published — the full INCI list with ordering retained, since concentration order carries meaning. Our capture rate is about 89%; some retailers publish partial lists or none.
We also detect formulation changes over time, which is how competitor reformulations become visible without waiting for press coverage. What we do not do is interpret the list for safety or efficacy — that is your regulatory and R&D function's work.
We parse and retain exactly what is published rather than forcing a standardised scale, because none exists across brands. Some use numeric codes, some descriptive names, some both.
Our parse rate is about 96.8%. For the deep-shade flag we derive the deeper end per brand and product line from the published range ordering rather than applying a fixed global rule, and the method is stated in the scope document so the flag is inspectable.
The same distinction that matters in quick commerce. A shade never ranged at a retailer is a category conversation; a ranged shade persistently unavailable is a supply conversation.
We keep them separate, so an availability percentage is computed against the shades actually ranged rather than against the full range the brand offers. Merging them makes a ranging decision look like a supply failure.
Yes, as structured mechanics rather than as offer text. GWP thresholds, sample and mini inclusion, bundle composition and validity windows.
These drive a large share of beauty promotion and are rarely tracked as mechanics, which means promotional intensity in this category is routinely understated by datasets that only track price.
Most beauty specialists do; general retailers and grocery beauty aisles frequently do not. We confirm per retailer during scoping.
Where shade availability is not published, we say so rather than substituting product-level stock. That single substitution would invert the availability conclusion, which is exactly the error the shade curve exists to prevent.
Yes — new product detection with first-seen dates, and separately shade_added_at for shades entering an existing range, which is how range extensions become visible.
Range extension into deeper shades is a particularly watched event in this category, and it is only detectable if shade-level history exists rather than product-level snapshots.
We quote individually. The main driver is whether shade-level collection is required — shade curves multiply record volume by the number of shades per product, which in foundation can be forty times product-level volume.
A defined category at product level sits at the lighter end; full catalogue at shade level with INCI capture and daily refresh sits higher. One scoping call, a free pilot on your own products within 48 hours, then a fixed monthly quote. Request a quote.
Send us a product list and retailers. We return shade curves with deep-shade flags, INCI capture and claims within 48 hours.
Free pilot, no card, no obligation. We'll confirm which retailers expose shade availability publicly.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Unlock retail insights with Noon Saudi Arabia Product Data Extraction to track prices, inventory, discounts, and product trends in real time.
Unlock Hertz & Avis Rental Car Data for Dynamic Pricing Intelligence to track rental rates, availability, and market trends in real time.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.