Core services
Enterprise Data Extraction

Scalable web, app and AI-powered collection across 40+ countries.

All 58 services →
New 2026
AI Training Data

Corpus building with provenance and opt-out compliance.

Learn more →
Free pilot
24-hour sample

We run collection on your own sources before you commit.

Get a sample →
58Services
40+Countries
DEVELOPER

Ready-Made Scrapers

Pre-built for top platforms. Self-serve, no setup.

View All →
TRY FREE

API Playground

Test endpoints instantly. No credit card.

Start Free →
28Tools
2SDKs
icons Delivery & SDKs
Streaming Crawl API Scheduler Realtime Alerts Webhook Delivery 🐍 Python SDK 💚 Node.js SDK
Need it managed instead?

Fixed monthly retainer, named engineer, no per-request metering.

Managed Data API →
HOT

Case Studies

How brands use Actowiz, with named outcomes.

Read →
FREE

Sample Datasets

Real output, no signup.

Download →
NEW

ROI Calculator

Model the return on a data engagement.

Calculate →
Platform · Shufersal

Shufersal & Israeli Grocery Data

You almost certainly do not need this scraped. Israeli law requires the prices to be published — the work is making thirty-five incompatible feeds agree.

Israel is the unusual case where the data is legally required to be published. Under the 2014 Food Act, supermarket chains with three or more stores must publish full price, promotion and store files publicly, updated through the day. So scraping Shufersal is the wrong approach. The real work is normalising roughly thirty-five chains that publish in incompatible formats — and that is what we would actually be selling you.

It is rare that the honest answer to 'can you scrape this' is 'you should not need to'. This is one of those cases, and pretending otherwise would be charging for the wrong thing.

Free pilot on your own Shufersal list, returned in 24 hours. No card, no trial clock — and you keep the sample data either way.

israel_normalised_2026-08-25.jsonl LIVE FEED
{"chain":"chain-a","store_id":"0412", "item_code":"7290000*** ","item_code_type":"barcode", "gtin_derived":"07290000***","gtin_confidence":0.99, "price":12.90,"currency":"ILS", "file_published_at":"2026-08-25T06:14:00+03:00", "chain_publication_cadence":"incremental_intraday + daily_full"} {"chain":"chain-b", "item_code":"INT-88120","item_code_type":"internal", "gtin_derived":"null","gtin_confidence":0.00, "caution":"internal code in the barcode field — a naive cross-chain join silently mismatches here"} {"chain":"platform-c", "own_sales_scope_note":"own dark-store range only; partner retailers on the same app are NOT in this file", "note":"a legal scope boundary, not a data gap"}
3 of 4,880,110 rows · 35 chains normalisedingested from LEGALLY PUBLISHED files · not scraped · schema v1.0

Independence and trademarks. Actowiz Solutions is not affiliated with, endorsed by or connected to Shufersal or its owners. Shufersal and related marks belong to their respective owners, used here only to name the publicly accessible source this service collects from.

Our Data Powers
B2C Marketplace
amazon
D2C + Marketplace
NYKAA
D2C + Marketplace
Walmart
FMCG Marketplace
udaan
Food Delivery
Uber Eats
Quick Commerce
blinkit
Taxi Aggregator
Uber
E-Commerce
Tmall
Shufersal at a glance

How we handle Shufersal specifically

Platform-specific handling, not a generic retail template pointed at a different domain.

The law
2014 Food Act, price transparency. In effect since 2015
Who it covers
Chains with three or more stores
What must be published
Full price, promotion and store files — publicly
Frequency
Updated through the day, with full daily snapshots per store
So scraping is
The wrong approach. The files are already public
The actual work
Normalising ~35 chains publishing incompatible formats
Format reality
Gzipped XML, ZIP-wrapped, UTF-16, different dialect per portal
Scope limit
The law covers each retailer's own sales, not partners on its platform
Platform specifics

Why the honest answer is not 'yes we can scrape it'

These are the reasons a Shufersal dataset needs its own handling rather than a shared retail schema.

The data is mandated, published and public

Israel's 2014 Food Act — the price transparency regulation — requires supermarket chains above a small size threshold to publish their price, promotion and store files publicly. The requirement came into effect in 2015 and it covers roughly thirty-five chains.

The files are published on chain portals, refreshed through the day, with full per-store snapshots daily.

What that means for a request to scrape Shufersal

The prices you want are already published. Building a scraper against the storefront would produce a worse dataset than the file the chain is legally required to put online, more slowly, and with more fragility.

We say this rather than quoting for the scraper, because quoting for it would be selling you the more expensive version of a thing you can have more cheaply.

What the honest engagement is

If you need Israeli grocery pricing at scale, the work is ingestion and normalisation across the published files — not extraction from web pages. That is a real piece of engineering and it is what we would scope.

Where the difficulty actually is

Mandated publication does not mean usable publication. The chains publish to the letter of the requirement and no further, and the practical result is a set of feeds that do not agree with each other.

  • Formats differ per portal. Gzipped XML, ZIP-wrapped files, and encodings including UTF-16 alongside UTF-8.
  • Portal behaviour differs. Some sign download URLs with a short expiry, so a listing must be followed immediately. Some are geo-restricted.
  • Identifiers are inconsistent. Item codes are usually barcodes, but chains also publish internal codes in the same field, so a cross-chain join needs a canonical identifier derived rather than taken.
  • Hebrew text arrives as published, in several encodings, and right-to-left handling has to survive the whole pipeline.
  • Update cadence differs — incremental files through the day, full snapshots daily, and not on the same schedule per chain.

Making thirty-five of those agree on one schema, with a derived canonical identifier and a stated freshness per chain, is the deliverable. It is a normalisation problem, not a collection one.

Two limits of the mandated data, worth knowing before you build on it

The law covers a retailer's own sales

Where a chain operates a platform carrying other retailers, the published file contains its own sales only. A delivery platform's file contains its own dark-store range, not the partner supermarkets listed on the same app.

That is a coverage boundary rather than a gap in the data, and it matters if you assume a platform's file represents everything sold through it.

Kosher certification

Certification is a significant purchase driver in this market and it appears in product data at varying levels of detail depending on the chain.

We capture certification text exactly as published, with verified_by_actowiz false. Whether a certification is current, and which certifying authority a buyer accepts, are matters for the issuing body and the buyer — not something extraction or normalisation can determine.

What we would not do

Interpret certification levels, rank them, or map one authority's certification onto another's. Those are religious and commercial determinations, and a vendor making them in a data field is overreaching.

Scope

What we collect on Shufersal, and what we do not

The right column matters more than the left. Anyone can list fields; the limits are what tell you whether the dataset will hold up.

✅ What we collect

  • Ingestion and normalisation across the published price files
  • One schema across chains, with the source chain on every record
  • A derived canonical identifier, since published item codes are inconsistent
  • Per-chain freshness stated, because publication cadence differs
  • Hebrew text retained exactly as published, encodings handled
  • Store-level records, since files are published per store
  • Promotion files alongside price files, kept as separate mechanics
  • Certification text as published, flagged unverified
  • The own-sales scope boundary stated per chain

❌ What we do not, and why

  • A scraper built against storefronts where the file is published
  • Certification levels interpreted, ranked or mapped between authorities
  • A canonical identifier assumed from a published item code without derivation
  • Partner-retailer sales read from a platform's own-sales file
  • Personal data of any kind

Core Shufersal fields

The full dictionary is agreed during scoping. These are the fields specific to this platform.

Field What it is on this platform
chain / store_id Source chain and store. Files are published per store
item_code / item_code_type As published, and what kind of code it is
gtin_derived / gtin_confidence Canonical identifier, derived rather than assumed
item_name_he Hebrew as published, encoding normalised
price / currency As published in the file
promo_mechanic / promo_source_file Promotions kept separate from price
file_published_at / ingested_at Freshness, which differs per chain
chain_publication_cadence Stated, since it is not uniform
own_sales_scope_note The legal scope boundary for this chain's file
certification_text / verified_by_actowiz As published, never interpreted
store_count_observed How many stores the batch actually covered
Use cases

What teams do with Shufersal data

Cross-chain Israeli price comparison

Thirty-five chains normalised to one schema with a derived canonical identifier, which is the whole difficulty and the whole value in this market.

Brand price monitoring across the market

A brand's products across every chain legally required to publish, which is close to complete market coverage rather than a sample.

Promotion mechanic tracking

Promotion files ingested alongside price files and kept as separate mechanics, rather than collapsed into an effective price.

Store-level price dispersion

Files are published per store, so intra-chain dispersion is directly observable rather than inferred.

The 24-hour sample — run on your sources, not ours

Send us a Shufersal item or category list. We run real collection against it and return the output within 24 hours, with the platform-specific fields populated so you can check them yourself rather than take our word for it.

  • Real extraction from your actual sources
  • Returned within 24 hours
  • Coverage and QA note included
  • You keep the data either way
  • No card, no trial clock
  • Named engineer on the call
Get my free sample Book a 20-min scoping call Reply within one business day. Reference calls available under NDA.
How we engage

Three ways to engage us

Same collection pipeline and QA underneath. The difference is who holds the schedule and how the data reaches you.

Managed service (most common)

We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.

  • Dedicated engineer assigned to your account
  • Site changes fixed by us, not reported to you
  • Scheduled delivery to your warehouse or S3
  • Named contact on Slack or email

Best fit: Teams who need the data, not the infrastructure.

API access

The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.

  • On-demand and scheduled endpoints
  • Rate limits agreed to your load profile
  • Sandbox keys for integration testing
  • Versioned schema with deprecation notice

Best fit: Product and engineering teams building on live data.

One-time or project extraction

A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.

  • Fixed scope agreed in writing upfront
  • Single delivery with full QA report
  • Methodology documented for your records
  • Converts to managed if you want continuity

Best fit: Research, strategy and diligence work with a deadline.

Pricing

Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.

Shufersal is usually collected alongside its competitors

Almost nobody buys a single platform in isolation. Shufersal data becomes useful when it sits next to the competitor set on one schema, refreshed on one schedule, so a price index or availability comparison is genuinely like-for-like.

That is what grocery data scraping covers, and a Shufersal-only engagement can be expanded into it without rebuilding. If you already know you need several platforms, start there instead — it is the same pipeline and usually the better scoping conversation.

FAQ

Shufersal data scraping: frequently asked questions

Platform-specific questions, including what cannot be collected here.

We could, and you should not pay us to. Israeli law requires chains with three or more stores to publish full price, promotion and store files publicly, updated through the day.

A scraper built against the storefront would produce a worse dataset than the file the chain is already required to publish, more slowly and with more fragility. We would rather tell you that than quote for it.

Ingestion and normalisation. The chains publish to the letter of the requirement and no further, so the feeds do not agree: different formats, different encodings including UTF-16, signed URLs with short expiry, geo-restricted portals, and inconsistent item codes.

Making roughly thirty-five of those agree on one schema with a derived canonical identifier and a stated freshness per chain is the work.

Because it is usually a barcode but not always — chains also publish internal codes in the same field. A cross-chain join built on it will silently mismatch.

We derive a canonical identifier with a confidence value rather than taking the published code at face value.

No. The law covers each retailer's own sales. Where a chain operates a platform carrying other retailers, the published file contains its own range — a delivery platform's file contains its dark-store range, not the partner supermarkets on the same app.

That is a scope boundary worth knowing before building on the assumption it is complete.

We capture certification text exactly as published, with verified_by_actowiz false. It appears at varying levels of detail depending on the chain.

What we will not do is interpret certification levels, rank them, or map one authority's onto another's. Those are religious and commercial determinations, and putting them in a data field would be overreaching.

We quote individually, and it is priced as a normalisation engagement rather than a collection one — chain count and refresh cadence are the drivers.

One scoping call, a free sample within 24 hours across a subset of chains, then a fixed monthly quote. If a public dataset already covers your need, we will point you at it. Request a quote.

See real Shufersal data before you commit to anything

Send us an item or category list. We return the output within 24 hours with the platform-specific fields populated.

Free pilot, no card, no obligation. If we cannot collect a field you need on this platform, the sample shows you that too.

Social Proof That Converts

Trusted by Global Leaders Across Q-Commerce, Travel, Retail, and FoodTech

Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.

4,000+ Enterprises Worldwide
50+ Countries Served
20+ Industries
Join 4,000+ companies growing with Actowiz →
Real Results from Real Clients

Hear It Directly from Our Clients

Watch how businesses like yours are using Actowiz data to drive growth.

1 min
★★★★★
"Actowiz Solutions offered exceptional support with transparency and guidance throughout. Anna and Saga made the process easy for a non-technical user like me. Great service, fair pricing!"
TG
Thomas Galido
Co-Founder / Head of Product at Upright Data Inc.
2 min
★★★★★
"Actowiz delivered impeccable results for our company. Their team ensured data accuracy and on-time delivery. The competitive intelligence completely transformed our pricing strategy."
II
Iulen Ibanez
CEO / Datacy.es
1:30
★★★★★
"What impressed me most was the speed — we went from requirement to production data in under 48 hours. The API integration was seamless and the support team is always responsive."
FC
Febbin Chacko
-Fin, Small Business Owner
icons 4.8/5 Average Rating
icons 50+ Video Testimonials
icons 92% Client Retention
icons 50+ Countries Served

Join 4,000+ Companies Growing with Actowiz

From Zomato to Expedia — see why global leaders trust us with their data.

Why Global Leaders Trust Actowiz

Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.

icons
7+
Years of Experience
Proven track record delivering enterprise-grade web scraping and data intelligence solutions.
icons
4,000+
Projects Delivered
Serving startups to Fortune 500 companies across 50+ countries worldwide.
icons
200+
In-House Experts
Dedicated engineers across scrapers, AI/ML models, APIs, and data quality assurance.
icons
9.2M
Automated Workflows
Running weekly across eCommerce, Quick Commerce, Travel, Real Estate, and Food industries.
icons
270+ TB
Data Transferred
Real-time and batch data scraping at massive scale, across industries globally.
icons
380M+
Pages Crawled Weekly
Scaled infrastructure for comprehensive global data coverage with 99% accuracy.

AI Solutions Engineered
for Your Needs

LLM-Powered Attribute Extraction: High-precision product matching using large language models for accurate data classification.
Advanced Computer Vision: Fine-grained object detection for precise product classification using text and image embeddings.
GPT-Based Analytics Layer: Natural language query-based reporting and visualization for business intelligence.
Human-in-the-Loop AI: Continuous feedback loop to improve AI model accuracy over time.
icons Product Matching icons Attribute Tagging icons Content Optimization icons Sentiment Analysis icons Prompt-Based Reporting

Connect the Dots Across
Your Retail Ecosystem

We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.

icons
Analytics Services
icons
Ad Tech
icons
Price Optimization
icons
Business Consulting
icons
System Integration
icons
Market Research
Become a Partner →

Popular Datasets — Ready to Download

Browse All Datasets →
icons
Amazon
eCommerce
Free 100 rows
icons
Zillow
Real Estate
Free 100 rows
icons
DoorDash
Food Delivery
Free 100 rows
icons
Walmart
Retail
Free 100 rows
icons
Booking.com
Travel
Free 100 rows
icons
Indeed
Jobs
Free 100 rows

Latest Insights & Resources

View All Resources →
thumb
Blog

How the US Grocery Price Inflation Tracker 2026 Helps Retailers Manage Rising Food Costs and Pricing Decisions

Track the US Grocery Price Inflation Tracker 2026 to monitor food price trends, category changes, and inflation insights for smarter decisions.

thumb
Case Study

How Brands Leverage KSA Hungerstation Menu Pricing Scraping API for Real-Time Food Delivery Intelligence

Explore KSA Hungerstation Menu Pricing Scraping API to track menu prices, competitor changes, and food delivery market trends in Saudi Arabia.

thumb
Report

Namshi Fashion & Beauty Data Intelligence

Namshi Fashion & Beauty Data Intelligence helps brands track prices, products, availability, assortment, and trends for smarter MENA market decisions.

Start Where It Makes Sense for You

Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.

icons
Enterprise
Book a Strategy Call
Custom solutions, dedicated support, volume pricing for large-scale needs.
icons
Growing Brand
Get Free Sample Data
Try before you buy — 500 rows of real data, delivered in 2 hours. No strings.
icons
Just Exploring
View Plans & Pricing
Transparent plans from $500/mo. Find the right fit for your budget and scale.
Get in Touch
Let's Talk About
Your Data Needs
Tell us what data you need — we'll scope it for free and share a sample within hours.
  • icons
    Free Sample in 2 HoursShare your requirement, get 500 rows of real data — no commitment.
  • icons
    Plans from $500/monthFlexible pricing for startups, growing brands, and enterprises.
  • icons
    US-Based SupportOffices in New York & California. Aligned with your timezone.
  • icons
    ISO 9001 & 27001 CertifiedEnterprise-grade security and quality standards.
Request Free Sample Data
Fill the form below — our team will reach out within 2 hours.
+1
Free 500-row sample · No credit card · Response within 2 hours

Request Free Sample Data

Our team will reach out within 2 hours with 500 rows of real data — no credit card required.

+1
Free 500-row sample · No credit card · Response within 2 hours