Rating metrics
The full shape, not just the average.
- Rating average and full distribution
- Review count and verified share
- Rating trend over trailing windows
- Distribution shift over time
- Rating by variant where separated
Aggregate signal and themes, not reviewer dossiers.
Review data is where scraping most easily drifts into building profiles of individuals. Everything commercially useful here is aggregate or thematic. Reviewer identity adds nothing to the analysis and a great deal to your risk.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Review and ratings data scraping is the automated collection of publicly posted review content and rating metrics across retail sites, marketplaces, app stores and service review platforms.
Almost everyone starts by asking for the rating average. It is the least informative field available.
Full distribution, verified-purchase share, review velocity against prior periods, rating trend over trailing windows, cross-platform gaps for the same item, and extracted themes with sentiment and trend direction. The average is included, but it is not where the value is.
Reviewer names, profile links, review histories and any reviewer-identifying information are not collected. Review text, rating, date and verified flag are — those describe a product. A reviewer's name plus their full review history across products is a profile of a person, and it adds nothing to product analysis while adding materially to your exposure.
Theme extraction is the highest-value output. Cross-platform comparison is the most commonly missed.
The full shape, not just the average.
The leading indicator.
What people actually say, structured.
The raw material for your own analysis.
Because a single platform is not the rating.
Handled with stated limits.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 90+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
item_key / platform |
string / enum | Cross-platform item identity and the platform the metrics came from | Every run |
rating_avg / rating_count |
decimal / int | Rating average and total review count | Daily |
distribution |
object | Full star distribution, which carries more information than the average | Daily |
verified_share_pct |
decimal | Share of reviews flagged as verified purchases where the platform indicates it | Daily |
velocity_30d / velocity_prev_30d |
int | Review volume in trailing windows, so momentum is visible not just level | Daily |
rating_trend_90d |
decimal | Rating movement over a trailing window, which the lifetime average conceals | Daily |
themes |
array | Extracted themes with mention counts, sentiment and trend direction | Weekly |
review_text |
array | Review text with rating, date and verified flag, without reviewer identity | Daily |
cross_platform_gap |
decimal | Rating difference for the same item against a reference platform | Daily |
anomaly_flags |
array | Unusual velocity or distribution patterns, flagged not labelled | Weekly |
reviewer_identity |
constant | Always not_collected, stated explicitly on every record | Every run |
reviewer_identity is a constant field reading not_collected. It exists so that anyone auditing the dataset can see the exclusion is structural rather than incidental.
Cross-platform coverage matters here more than depth on any one platform, because ratings differ systematically between them.
Some platforms publish reviews only to logged-in users or limit historical depth. We collect what is publicly accessible and state per platform how far back coverage extends rather than implying full history. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United States & United Kingdom | The deepest public review volumes across retail and app platforms, which makes theme extraction most reliable. |
| Germany, France & Netherlands | High review engagement with strong retailer review programmes and multilingual theme extraction demand. |
| India & GCC | Rapidly growing review volumes with high cross-platform variance for the same products. |
| Australia & Canada | Smaller review populations where distribution shape matters more because averages move on lower volume. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Product and quality teams dominate, with category, CX and insight teams close behind.
Product failure modes surface in reviews months before they appear in returns or warranty data.
Theme extraction with sentiment and trend across your products and competitors, plus velocity and distribution shift detection.
Return rate
Range decisions need to know which products are genuinely well received rather than which have high averages.
Full distribution, verified share and theme analysis by product and category, with cross-platform gaps included.
Category NPS
Voice-of-customer programmes rely on surveys with low response rates while reviews sit uncollected.
Structured theme data with sentiment and trend across products and platforms, refreshed continuously.
Insight coverage
Review performance on retailer listings affects conversion, and it differs by retailer without anyone tracking why.
Per-platform rating metrics and themes for your products, revealing where listings underperform and on what dimension.
Conversion rate
Public review sentiment is the visible face of service quality and is monitored inconsistently.
Rating trend, velocity and theme tracking across public review platforms with anomaly flagging.
Review sentiment trend
Review velocity and sentiment trend are observable leading indicators of product and brand trajectory.
Longitudinal review velocity, distribution and theme sentiment panels by brand and category.
Signal lead time
Four patterns, with the outcome each is judged on.
Themes are extracted from review text with sentiment and trend direction, so a rising negative theme is surfaced while the lifetime rating average is still unaffected.
Outcome: Product issues identified from review themes ahead of returns and warranty data.
Full star distributions, trailing rating trends and verified-purchase share are delivered, revealing polarised products and recent deterioration that lifetime averages conceal for months.
Outcome: Product assessment based on distribution shape and recent trend rather than an anchored average.
The same item is matched across platforms with rating gaps and theme differences computed, since populations and prompting differ systematically between platforms.
Outcome: Listing performance diagnosed per platform instead of assumed uniform.
Themes and sentiment are compared across competitor products in the same category, showing which attributes drive satisfaction and which competitors are criticised for what.
Outcome: Product development priorities informed by comparative review themes.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Quality monitoring relied on warranty and returns reporting, which lagged the customer experience by a full quarter.
Theme extraction with sentiment and trend across the product range, surfacing rising negative themes while lifetime averages were still unaffected.
The failure mode was identified from review themes well ahead of the returns signal.
Category buying compared products on rating average, treating a polarised 4.3 and a uniformly mediocre 4.3 as equivalent.
Full star distributions, verified-purchase share, trailing rating trend and theme sentiment across the category.
Range reviews began separating polarised products from consistently average ones, changing several delist decisions.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Theme extraction and cross-platform item matching are the work; collecting the average is trivial.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Clients frequently ask us to identify fake reviews. We provide anomaly signals and we deliberately stop short of labelling reviews as fake, and the distinction is not evasion — it is the limit of what public data supports.
Every one of those patterns has legitimate explanations. A velocity spike can be a promotion, press coverage or a review-request campaign. Low verified share can reflect a platform's flagging behaviour rather than fraud. Text similarity can come from prompted reviews with suggested wording. And detecting coordinated inauthentic behaviour reliably requires reviewer-level and account-level data, which is exactly what we do not collect.
Anomaly flags with the specific pattern named, so your team can investigate. We will not label a review or a product as fraudulent, because a false accusation is a serious matter and public data does not support the conclusion.
Platforms themselves have account-level signals we do not and should not have. If fraud is your primary concern, platform reporting mechanisms and specialist fraud vendors with platform relationships are better routes, and we will say so.
It is technically straightforward to collect reviewer names, profile links and full review histories. Most review scraping does. We do not, and the reasoning is worth setting out because it is a decision rather than a limitation.
Almost nothing for product analysis. Themes, sentiment, distribution, velocity and verified share answer the product questions. Knowing that a particular named person wrote a review does not improve a failure mode analysis or a category comparison.
Every record carries a reviewer_identity field with
the constant value not_collected. It is redundant
data by design, and it exists so that anyone auditing the dataset
— your DPO, a customer, a regulator — can see the
exclusion is structural rather than a claim on a webpage.
The same boundary applies across our social media, app store and local business services.
Platforms, item set and whether theme extraction is required are scoped first, since theme work is the labour-intensive component.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect publicly accessible review content and rating metrics. Reviewer names, profile links, review histories and any reviewer-identifying information are categorically excluded, and every record carries a reviewer_identity field stating not_collected. We do not label reviews as fraudulent, and we do not access platform accounts or authenticated review surfaces.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What product, quality and insight teams ask during evaluation.
No, categorically. Review text, rating, date and verified flag
are collected; reviewer names, profile links and review
histories are not. Every record carries a
reviewer_identity field with the constant value
not_collected so the exclusion is auditable.
The reasoning is practical as much as ethical: reviewer identity adds almost nothing to product analysis while creating personal data at scale that attracts subject access and deletion obligations the moment it lands in your systems.
We flag anomalies and we will not label reviews as fake. We surface unusual velocity, distribution anomalies, verified-share anomalies, sudden rating shifts and text homogeneity, naming the specific pattern.
Every one of those patterns has legitimate explanations — promotions, press coverage, prompted review campaigns, platform flagging behaviour. Reliable detection of coordinated inauthentic behaviour needs account-level data we deliberately do not collect. A false accusation is serious, so we deliver signals for your team to investigate rather than conclusions.
Because it conceals the two things that matter. Distribution shape distinguishes a product with a specific failure mode from a merely mediocre one at the same average. And a product with thousands of reviews barely moves its average even when recent reviews turn negative, so deterioration is invisible for months.
We deliver full distribution, trailing rating trend and review velocity against prior periods. The average is included; it is just not where the signal is.
About 89% on English reviews, lower on other languages and on very short reviews. Themes come with mention counts, sentiment and trend direction rather than a single sentiment score.
We publish the rate and report per-language performance rather than a single headline figure, because extraction quality varies materially and a team building product priorities on themes should know where the data is weaker.
Because review populations and prompting differ systematically. A retailer emailing purchasers gets a different sample than a marketplace where reviewing is unprompted, and app store review prompts differ again. Gaps of 0.3 to 0.5 stars for the same product are common.
We match items across platforms and compute the gap explicitly, since a single-platform average is not the product's rating. Theme differences by platform are often more informative than the rating gap itself.
As far back as each platform publicly exposes it, which varies considerably. Some platforms show full history; others paginate to a limit or surface only recent and highest-voted reviews.
We state per platform how far back coverage extends rather than implying full history. For products where deep history matters, we report what is retrievable during the pilot so you can judge before committing.
Where the platform separates them, yes. Many platforms pool reviews across a variation family, which means a review of a different colour or size appears on the item you are analysing.
We capture variant references where shown and flag where reviews are pooled, because pooled reviews on a poorly performing variant can drag an otherwise strong product — and that is a listing structure problem rather than a product problem.
We collect publicly posted review content without accounts or credentials. Review text about a product is product feedback; the sensitivity comes from reviewer identity, which we exclude categorically.
Platform terms often restrict automated access and we state that plainly. You receive a written methodology document per platform and a DPA before signature. In our experience the reviewer identity exclusion is what makes this category straightforward for a DPO to approve.
We quote individually. The drivers are platform count, item count, whether theme extraction is required, and refresh frequency. Theme extraction is the labour-intensive component and the main cost differentiator.
Rating metrics and velocity across a defined item set at daily refresh sits at the lighter end. Multi-platform coverage with theme extraction across large catalogues sits higher. One scoping call, a free pilot on your own items within 48 hours, then a fixed monthly quote. Request a quote.
Send us an item list and platforms. We return distributions, velocity, trends and extracted themes within 48 hours.
Free pilot, no card, no obligation. Reviewer identity is never included, on any engagement.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
A national price average is the arithmetic mean of your best and worst markets. Why geo-resolved price collection changes the numbers, and how to do it correctly.
A one-time extraction of up to 50,000 Wegmans products with pricing and nutrition attributes. Why single-location scoping and attribute completeness decide whether a bulk catalogue is usable.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.