Typed, unit-normalised attributes
So numeric and range constraints resolve.
- Values typed, not strings
- Units normalised to a stated base
- Ranges represented as min and max, not collapsed
- Parse basis on every derived value
Three companies asked us for this last quarter, all building natural-language product discovery. The feed they need is not the catalogue feed with a different label on it.
Three companies asked us for this last quarter, all building natural-language product discovery. The feed they need is not the catalogue feed with a different label on it.
Free pilot on your own sources, returned in 24 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
A person reading a product page fills gaps automatically. They see no mention of a feature and conclude nothing — they check elsewhere, or they ask. An agent does not do that unless the data tells it to.
dishwasher_safe is absent and the agent filters
for dishwasher safe, the product is excluded — and the shopper never learns it might have qualified.Every attribute as {value, basis} where it is known, and {value: null, reason} where it is
not. The reason distinguishes not published, published but unparseable, and not applicable to
this category — three situations an agent should handle differently.
Plus observed_at, freshness_minutes and a safe_to_recommend_until horizon
agreed with you, so the agent can degrade gracefully rather than asserting a stale price.
All of it derived from the same collection, shaped for a different consumer.
So numeric and range constraints resolve.
The single most important field family here.
So a stale recommendation is preventable.
So an agent's memory of a product survives.
Because the price alone is rarely the offer.
So an agent does not repeat a claim as fact.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
One record per product, shaped so a machine can act without guessing.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
product_key |
string | Persistent across runs, so agent memory survives | Every record |
title / brand / category |
string | As published, plus normalised category | Every record |
attributes |
object | Each as {value, basis} or {value: null, reason} | Every record |
attribute_fill_rate |
object | Per-attribute fill for this category, so gaps are known | Per batch |
price / currency / market |
number / string | Offer basics with market as a dimension | Every record |
in_stock / availability_reason |
boolean / string | Availability with a reason where false | Every record |
observed_at / freshness_minutes |
timestamp / number | How old the record is at delivery | Every record |
safe_to_recommend_until |
timestamp | The horizon you agreed, so the agent can degrade gracefully | Every record |
variant_axes |
array | Explicit rather than flattened into the title | Where applicable |
match_confidence |
number | On cross-retailer links, so weak matches can be excluded | Matched records |
verified_by_actowiz |
constant | False on claim fields. A claim is never a fact | Claim fields |
The per-attribute fill rate ships with every batch. An agent deployed on an attribute that populates on 30% of a category will fail in a way that looks like a model problem and is not.
Attribute density decides this, and it varies enormously by category.
We report the per-attribute fill rate on your own categories in the sample. A category where your key constraint attribute populates thinly is one where an agent will disappoint users, and it is better to know that before launch. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Mostly teams shipping a product, not teams running an analysis.
Building natural-language product search and finding that catalogue feeds cannot answer constraints.
Typed, unit-normalised attributes with explicit absence, so constraints resolve or the agent asks.
Answer quality
Needs to know which attributes are dense enough to build on before shipping a feature.
Per-attribute fill rate by category, delivered with every batch.
Feature viability
Wants agent-driven discovery without recommending out-of-stock or mispriced items.
Freshness in minutes and an agreed safe-to-recommend horizon per record.
Recommendation reliability
Wants their products to be findable by agents, which needs machine-readable attributes.
Attribute coverage audit against competitors, showing where a catalogue is invisible to agents.
Agent discoverability
Needs the same collection to serve both a warehouse and an agent without two pipelines.
One collection, two shapes — analytical records and agent records from the same source.
Pipeline economy
An agent asserting an unverified manufacturer claim is a liability.
Claims marked as claims, no inferred attributes, source per attribute.
Claim exposure
Four patterns, and one of them is defensive.
Typed and unit-normalised attributes so 'under 500 grams', 'fits a 27-inch monitor' and 'dishwasher safe' resolve against real values rather than being matched against marketing copy.
Outcome: Answers that hold up when a user checks them.
Explicit absence with a reason, so an agent can say it does not know instead of asserting a negative nobody published.
Outcome: The failure mode that damages trust fastest, removed.
Freshness in minutes plus an agreed horizon, so an agent degrades to 'check availability' rather than confidently recommending something out of stock.
Outcome: Merchant relationships that survive the agent.
Attribute coverage on your own products against competitors, showing where your catalogue cannot answer the constraints buyers will ask agents about.
Outcome: A fixable content gap identified before it costs share.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
If your catalogue already has typed attributes with explicit nulls, reshape it. Most do not.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 24 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
There are two ways an attribute can be not-true, and an agent must handle them differently.
| The maker said no | Nobody said | |
|---|---|---|
| Our field | {"value": false, "basis": "stated_by_manufacturer"} |
{"value": null, "reason": "not_published"} |
| Agent should | Exclude confidently | Include with a caveat, or ask |
| If collapsed | The agent excludes products that might qualify, and tells the user they do not — asserting something nobody claimed | |
Almost every product feed in the market collapses these, because a dashboard user reading a blank cell knows what it means and an agent does not.
Three reasons are distinguished, not one: not_published (the source is silent), unparseable (it was stated but not in a form we could type confidently), and not_applicable (the attribute is meaningless for this category). An agent can reasonably surface the second to a user and should ignore the third.
This is a data feed, not an agent, and the boundary is worth stating because the two are often conflated.
The moment we generate an attribute, your agent's provenance chain breaks. You could no longer tell a user — or a regulator, or a merchant — which facts came from the manufacturer and which came from a model.
That distinction is the thing agents will be judged on, and it is cheap to preserve now and expensive to reconstruct later.
The sample exists mainly to show you your real attribute coverage.
Not the fields — the questions. 'Under 500g', 'fits a 27-inch monitor', 'suitable for sensitive skin'. Those determine which attributes matter.
Before quoting. A key constraint attribute that populates on 30% of a category is a launch problem, and you should see it now.
Real records from your own catalogue with typed attributes, reasoned nulls and freshness stamps, so you can run your agent against it before committing.
How stale is too stale depends on your category and your promise. We set it per field group rather than one number for everything.
Live in 5 to 10 business days, with the per-attribute fill rate in every batch so coverage degradation is visible before users notice it.
JSON or JSONL for agent consumption, Parquet for bulk loads, or a REST endpoint with per-product lookup. Delivered to Amazon S3, Google Cloud Storage, Azure Blob, Snowflake, BigQuery or an endpoint you specify.
We collect only publicly accessible information. Manufacturer claims are captured as claims with verified_by_actowiz set false, and no attribute is inferred or modelled into an observation field.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 24 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
Straight answers for teams building on this.
Same collection, different shape. The catalogue feed is built for analysts and dashboards; this is built for a machine that cannot fill gaps by judgement.
The differences that matter: attributes are typed and unit-normalised, absence is explicit with a reason, and freshness is stamped so the agent can judge staleness. Most catalogue feeds fail all three.
Because an agent that cannot tell 'the maker says no' from 'nobody said' will exclude products that might qualify and tell the user they do not — asserting something nobody claimed.
That is the failure mode that destroys trust in an agent fastest, and it is entirely a data problem rather than a model problem.
No, deliberately. Embeddings should be generated against your model on your schedule, and a vendor's vectors go stale the moment you change models.
We deliver structured fields. The vector layer, the ranking and the prose generation are your product.
We can, and we will not put the result in an observation field. If you want inferred attributes, run the model on our data and keep the output in a column that is visibly a prediction.
The moment we generate an attribute, your provenance chain breaks and you can no longer tell a user or a merchant which facts came from the manufacturer.
It depends on your promise, which is why we set safe_to_recommend_until per field group rather than one number. Price and availability usually need a much shorter horizon than attributes.
The point is that the agent can see how old a record is and degrade to 'check availability' rather than recommending confidently from stale data.
You will see it in the sample, per attribute, before committing. Fashion and toys are typically thinner than electronics and industrial.
Thin coverage is not automatically fatal — an agent that asks a clarifying question is better than one that guesses — but it changes what you build, and it should change it before launch rather than after.
Yes, and for a discovery agent it usually must. Cross-retailer records arrive with a match confidence so weak links can be excluded by your own threshold rather than ours.
For a brand-side agent the same data doubles as a discoverability audit — showing where your catalogue cannot answer constraints that competitors can.
Yes, and it is where most of the demand we see is coming from. The additional considerations are pincode-level availability, which our quick-commerce work already handles, and multilingual titles, where we retain the original and do not machine-translate into the observation field.
We quote individually. Drivers are catalogue size, attribute breadth, source count and refresh frequency — refresh matters most because a short safe-to-recommend horizon means more observations.
One scoping call, a free sample within 24 hours with per-attribute fill rates, then a fixed monthly quote. Request a quote.
Tell us the constraints it must answer. We return real records with typed attributes, reasoned nulls and per-attribute fill rates.
No sales sequence. If your key constraint attribute is thin in your category, the sample is where you find that out rather than after launch.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
A national price average is the arithmetic mean of your best and worst markets. Why geo-resolved price collection changes the numbers, and how to do it correctly.
A one-time extraction of up to 50,000 Wegmans products with pricing and nutrition attributes. Why single-location scoping and attribute completeness decide whether a bulk catalogue is usable.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.