How Actowiz Solutions collected review & ratings data across Flipkart, Nykaa, Purplle, BigBasket, Blinkit & Zepto — structured, deduplicated, sentiment-ready.
A consumer brand operating across India's beauty, personal-care, and packaged-goods categories — selling through the full spread of platforms Indian shoppers actually use: Flipkart and Nykaa and Purplle for considered purchases, BigBasket and Blinkit and Zepto for the everyday and the instant. Their products lived on six platforms; their understanding of what customers thought of those products lived nowhere, in aggregate. Each platform showed reviews on its own pages, in its own format, and nobody on the brand's side had a unified, structured view of the voice of their customer across the places that voice was actually being expressed.
They came to Actowiz Solutions for one thing: all the review and ratings data for their products (and key competitors') across those six platforms, in one clean structured dataset — deduplicated, normalised, and ready for the sentiment and theme analysis their insights team wanted to run.
Review data across six diverse platforms is a harder problem than it appears:
Per-platform extraction feeding one normalised schema: product reference, platform, rating, review title and body, review date, verified-purchase flag, helpful votes, and platform-specific enrichment fields (beauty tags from Nykaa/Purplle, delivery-experience signals from q-commerce) preserved in typed extensions — so the common structure enables cross-platform analysis while the platform-specific richness isn't flattened away.
Language identification per review (including Hinglish and code-mixed tagging), with processing that preserves meaning across English, Hindi, Hinglish, and regional languages — the multilingual depth from our regional-language work, applied to the review-sentiment use case. Original text retained; language tagged for downstream analysis.
Full review collection across pagination (not just first-page or "top" reviews), with deduplication within and across platforms, so the dataset represents the genuine distribution of customer sentiment rather than a skewed sample.
Reviewer identity masked during collection — names, profile references, and incidental personal details in text handled per DPDP — retaining the commercial signal (rating, text themes, sentiment, verified flag) and nothing that creates personal-data liability. This is both compliant and sufficient: the brand needs the voice, not the identity.
Incentivised-review patterns, template spam, and AI-generated text flagged and filtered as a curation stage, so the delivered dataset is sentiment-analysis-ready rather than polluted.
Reviews delivered structured for the client's analysis — clean text, language tags, ratings, and platform context — with an optional theme-and-sentiment enrichment layer (aspect-level: product quality, delivery, value, packaging) that decomposes the voice into the dimensions the insights team wanted rather than a single star average.
Delivered in the client's structured format on a recurring cadence with historical retention (so sentiment trends over time became visible), per-record lineage, and DPDP-mapped compliance documentation.
| Field | Value* |
|---|---|
| Product | Sample Face Serum |
| Platform | Nykaa |
| Rating | 4 / 5 |
| Language | Hinglish |
| Review (masked) | "Texture bahut light hai, absorbs fast, but glass dropper feels fragile" |
| Beauty tags | Skin: combination; Concern: dullness |
| Verified | Yes |
| Reviewer identity | [masked at edge] |
| Platform | Reviews* | Avg Rating* | Top Positive Theme* | Top Negative Theme* |
|---|---|---|---|---|
| Flipkart | 2,140 | 4.2 | Value | Packaging |
| Nykaa | 1,880 | 4.4 | Texture/results | Dropper fragility |
| Blinkit | 640 | 4.1 | Fast delivery | — |
| Zepto | 510 | 4.0 | Availability | Occasional stock issues |
Sample data — illustrative of deliverable format. Actual delivery is review-level with language tags, enrichment fields, and masked identity.
| Metric | Value* |
|---|---|
| Platforms | 6 (Flipkart, Nykaa, Purplle, BigBasket, Blinkit, Zepto) |
| Reviews collected (brand + competitors) | Hundreds of thousands |
| Languages handled | English, Hindi, Hinglish, regional |
| PII masking | 100% reviewer identity masked at edge |
| Noise filtered (incentivised/AI/spam) | Flagged & removed as curation stage |
| Delivery | Structured, recurring, sentiment-ready |
| Time to first delivery | 3 weeks |
Representative engagement figures — illustrative of project structure.
The brand's insights team got, for the first time, a single unified view of their customer's voice across every platform their products sold on — and the cross-platform view immediately surfaced things single-platform browsing had hidden. The same serum praised for results on Nykaa was criticised for packaging on Flipkart and delivery on q-commerce — three different truths about one product, actionable by three different teams (formulation held steady, packaging flagged for redesign, q-commerce fulfilment raised with partners). The aspect-level sentiment decomposition meant "4.2 stars" became "loved for texture, dinged for a fragile dropper" — a product-development instruction rather than a vanity metric.
The multilingual handling proved its worth in the Indian context specifically: a large share of the most detailed, most useful reviews were Hinglish and code-mixed, and a pipeline that mishandled them would have discarded exactly the richest signal. The q-commerce reviews, distinct in character, gave the brand its first structured read on delivery-and-freshness perception — the axis that increasingly decides repeat purchase in instant commerce.
The engagement continues as a recurring feed, with sentiment trends now tracked over time (did the packaging redesign move the packaging-complaint rate?) and the panel expanding to more products and competitors — the voice-of-customer becoming a monitored metric rather than an occasional manual audit.
Every multi-platform consumer brand faces the same blind spot: their customer's voice is fragmented across the platforms they sell on, in multiple languages, in incompatible formats, laced with PII and noise. The transferable design: a unified-but-extensible review schema across platforms, genuine multilingual and code-mixed handling, complete deduplicated collection, PII masking at the edge, noise filtering, and aspect-level sentiment structuring. The unified, clean, compliant voice-of-customer is the asset — and in India specifically, the multilingual layer is what makes it real.
Yes — a unified schema captures common review fields across all six platforms while preserving platform-specific richness (beauty tags from Nykaa/Purplle, delivery signals from q-commerce) in typed extensions, enabling cross-platform analysis without flattening.
With language identification (including Hinglish/code-mixed tagging) and meaning-preserving processing — essential in India, where a large share of the most detailed reviews are code-mixed and would otherwise be lost as noise.
No — reviewer identity is masked at the edge per DPDP, retaining only the commercial signal (rating, text, themes, verified flag). The brand needs the voice, not the identity.
Yes — an optional enrichment layer decomposes reviews into aspect-level sentiment (quality, delivery, value, packaging), turning star averages into actionable product insight. Contact Actowiz Solutions to scope a review-data programme.
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Tyres Categories data collection from Lazada and Tuhu App helps businesses track tyre prices, brands, availability, and assortment for market insights.
Track Scrape Ready-to-Cook Cut Veg Product Data from Blinkit TN to monitor prices, availability, SKUs, and trends for smarter retail insights.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.