Organic results
Positions with the context that makes them meaningful.
- Position, URL, domain and title
- Pixel depth on the rendered page
- Above-fold flag by device
- Displayed snippet text
- Sitelink and rich result presence
With location, device and AI Overview citations captured, not a single global rank.
There is no such thing as your ranking. There is your position for a keyword, in a location, on a device, in a language, at a moment. Any vendor reporting one number has silently chosen the other four for you.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
SERP data scraping is the automated collection of search engine results pages: the organic results and their order, sponsored placements, and the growing set of SERP features that occupy the page — AI Overviews, featured snippets, local packs, shopping carousels, video blocks and question boxes.
The central point about this category is that a search result is a function of four variables, and collapsing them produces a number that describes nobody's experience.
Being position one means less than it used to, because the page above position one has grown. On many commercial queries a mobile SERP now carries an AI Overview, four sponsored slots and a shopping pack before the first organic result appears.
So we capture pixel depth — how far down the page a
result actually sits — alongside position, plus an
organic_above_fold flag. A position-one result 1,800
pixels down the page is not a position-one experience, and no
amount of rank reporting will show that.
Where an AI-generated overview appears, the commercially important question has shifted from "do I rank" to "am I cited". We capture AI Overview presence, the domains it cites, citation count and its position on the page — per keyword and location, since AI Overview presence itself varies by both.
Search volume, click data or clickstream. Volume figures come from search engine tools and licensed data providers; click behaviour is not observable from a SERP. Anyone offering scraped search volume is passing off a model or a licensed source.
Feature detection and AI Overview citation tracking are the fastest-growing use cases.
Positions with the context that makes them meaningful.
The new visibility layer.
How much of the page is bought.
Where local intent is served.
The full composition of the page.
Handled honestly.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 120+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
keyword / engine |
string / enum | Query and search engine, since results differ substantially between engines | Every run |
country / location / device / language |
string / enum | The four dimensions that define a result, all mandatory | Every run |
observed_at |
timestamp | Exact capture time, essential because SERPs change within the day | Every run |
features_present |
array | Every SERP feature detected, classified into a standard taxonomy | Every run |
ai_overview |
object | Presence, cited domains, citation count and position on page | Every run |
organic |
array | Positions with URL, domain, title, snippet and pixel depth | Every run |
pixel_depth |
int | Vertical position on the rendered page, which position number alone does not convey | Every run |
organic_above_fold |
boolean | Whether any organic result appears above the fold on that device | Every run |
paid_slots_top / paid_slots_bottom |
int | Sponsored placement counts, so bought share of the page is measurable | Every run |
local_pack |
array | Local pack members with position and displayed rating where present | Every run |
sample_variance |
decimal | Where repeat sampling is configured, observed variance for the same query | Per configured runs |
SERPs are not deterministic. Two requests seconds apart can differ, and we do not pretend otherwise. Where accuracy matters, repeat sampling is configurable and observed variance is delivered rather than hidden behind a single reading.
Location granularity is the main cost lever. City-level is usually sufficient; sub-city is available where local intent is strong.
We collect what is publicly returned to an ordinary search request. We do not use search engine advertising accounts, keyword planner data or any authenticated surface, and we do not provide search volume. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United States & United Kingdom | Where AI Overviews rolled out earliest and most broadly, making citation tracking most urgent for brands here. |
| Germany, France & Spain | Multilingual SERPs where language and location interact, so single-dimension rank tracking fails most visibly. |
| India & GCC | High-growth search markets with heavy mobile skew, where device-specific feature composition differs sharply from desktop. |
| Australia & Canada | Strong local intent patterns making sub-city location targeting particularly valuable for multi-location businesses. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
SEO teams and agencies dominate; AI search visibility tracking is the fastest-growing reason.
Rank reporting shows positions without showing how much of the page sits above them, so improving position stops correlating with traffic.
Positions with pixel depth and above-fold flags, plus full feature inventory, so visibility is measured as page reality rather than ordinal rank.
Organic sessions
AI Overviews answer queries without a click, and you cannot tell whether you are cited or invisible.
AI Overview presence with cited domains and counts per keyword and location, tracked over time.
AI citation share
Client reporting needs defensible rank and feature data across many locations and devices.
Consistent multi-location, multi-device SERP panels with feature detection, delivered on a reporting schedule.
Client retention
Local pack composition differs by neighbourhood, and national rank tracking cannot see it.
Sub-city location targeting with local pack membership and position captured per location.
Local pack presence
Shopping packs and sponsored slots consume the page, and organic reporting ignores them entirely.
Paid slot counts, shopping pack participants and paid share of above-fold space alongside organic position.
Paid and organic share
Studying search result composition and information sources requires reproducible, documented SERP capture.
Documented SERP capture with all four dimensions recorded and sampling variance reported.
Reproducibility
Four patterns, with the outcome each is judged on.
AI Overview presence and cited domains are captured per keyword and location, so citation share can be measured over time and against competitors, including where an overview replaces organic clicks entirely.
Outcome: AI search visibility measured rather than assumed, with citation gaps identified by keyword cluster.
Positions are captured with pixel depth and above-fold flags per device, so a position-one result sitting below an AI Overview, four ads and a shopping pack is correctly identified as low visibility.
Outcome: Visibility reporting that tracks traffic reality instead of ordinal rank.
SERPs are collected per targeted location including sub-city where local intent is strong, capturing local pack membership and position per location.
Outcome: Local visibility managed per catchment rather than as a national average.
Full feature inventory plus paid slot counts and organic positions are combined into a share-of-page view by keyword, showing who occupies what and how it changes.
Outcome: Competitive share measured across the whole page including paid and features, not organic alone.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Reporting showed positions rising across a commercial keyword set, but organic sessions declined, and nobody could reconcile the two.
Collection with pixel depth, above-fold flags and full feature inventory including AI Overview presence, per device and location.
Position-one results were sitting below an AI Overview, four ads and a shopping pack on mobile, which explained the divergence.
The publisher suspected AI-generated answers were reducing clicks but had no measurement of whether it was being cited or bypassed.
AI Overview presence and cited domain capture across the topic keyword set, tracked by location and over time.
Citation share became measurable by topic cluster, redirecting content effort to where citations were being lost.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Feature detection changes constantly as search engines change layouts, which makes this permanent maintenance rather than a build.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Rank tracking tools present a single number per keyword per day. This implies determinism that does not exist, and the implication causes real analytical damage.
Where accuracy matters, repeat sampling is configurable: the same
query is collected multiple times within a window, and we deliver
observed variance in a sample_variance field
alongside the reading. That turns a false certainty into a
measured confidence interval.
Every record also carries a precise
observed_at timestamp, so a position can be tied to a
moment rather than to a day. For keyword sets tracked through an
algorithm update, this is the difference between seeing a change
and guessing at one.
We would rather deliver a range with variance than a single number that looks authoritative and is not reproducible. Teams making content investment decisions on rank movements of one or two positions are frequently reacting to noise, and only variance data reveals that.
For a growing share of informational and commercial queries, an AI-generated overview now sits above the organic results and answers the query directly. This changes what SERP data needs to capture.
AI Overview presence, cited domains, citation count and page position are captured per keyword, location and device on every run. Because presence itself is variable, repeat sampling is particularly valuable here — a single reading can miss an overview that appears on most requests.
What we do not do is claim to explain why a domain is cited. Citation selection is not documented by search engines, and vendors offering optimisation formulas for it are guessing. What we provide is accurate observation of who is cited, for which queries, where, and how that changes — which is the factual basis your content strategy needs. Our news data service is often paired with this for tracking which publishers gain citation share.
Keyword set, location granularity, devices and sampling configuration are scoped first, since those four multiply volume.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect what is publicly returned to an ordinary search request, at controlled request rates. We do not use search engine advertising accounts, keyword planner data, or any authenticated surface, and we do not provide search volume or click data. Search engine terms restrict automated querying and we state that plainly; methodology is documented per engine.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What SEO, content and agency teams ask during evaluation.
Yes — presence, cited domains, citation count and position on the page, per keyword, location and device. This is the fastest-growing reason clients come to us.
What we will not do is claim to explain why a domain gets cited. Citation selection is not documented by search engines, and formulas for it are guesswork. We provide accurate observation of who is cited for which queries and how it changes, which is the factual basis a content strategy needs.
Because a search result does not exist independently of them. Results differ by city and sometimes neighbourhood, and mobile and desktop SERPs differ in feature composition and in how much page organic results occupy.
A tool reporting one rank number has silently chosen a location and device for you. We make all four dimensions — keyword, location, device, language — mandatory fields, so you always know what the number describes.
How far down the rendered page a result actually sits, in pixels. Position one used to mean top of page; on many mobile commercial queries it now sits below an AI Overview, four sponsored slots and a shopping pack.
We deliver pixel depth and an
organic_above_fold flag alongside position,
because a position-one result 1,800 pixels down is not a
position-one experience. This is usually the field that
explains why improving rank stopped improving traffic.
No. Search volume is not published in search results — it comes from search engine advertising tools and licensed data providers. Anyone offering scraped search volume is passing off a model or a licensed source.
We collect what is observable on the page. For volume, use the engine's own tools or a licensed provider, and join it to our data on the keyword field.
Not perfectly, and we do not pretend otherwise. The same query seconds apart can differ due to ongoing experimentation, data centre index variation, location inference and feature triggering thresholds.
Where accuracy matters, repeat sampling is configurable and we
deliver observed variance in a
sample_variance field. That converts false
certainty into a measured confidence range — and it
often reveals that a team reacting to a one-position movement
is reacting to noise.
Country, city, and sub-city where local intent is strong. Sub-city granularity matters for local pack work, since pack composition can differ between neighbourhoods in the same city.
Location count is one of the main cost multipliers, so we scope it deliberately. For most national SEO work, a handful of representative cities answers the question; for multi-location businesses, per-location collection is the point.
Google as the primary, plus Bing, DuckDuckGo, Yandex, Baidu, Naver and Yahoo Japan where relevant to your markets. We also collect Amazon, YouTube and app store search results, which behave as their own search ecosystems.
Feature detection is most complete on Google because it has the richest feature set; on other engines we detect what exists. We state per engine what feature coverage looks like during scoping.
Search engine terms restrict automated querying, and we say that rather than glossing over it. Our practice is to collect what is publicly returned to an ordinary request, at controlled rates, without advertising accounts or any authenticated surface.
You receive a written methodology document per engine and a DPA before signature. Notably, rank tracking is an established industry practice used by essentially every SEO team and tool, which is useful context for that conversation — though it is not a legal conclusion, and your counsel should form their own.
We quote individually, and here the quote is driven by four multipliers: keywords, locations, devices and refresh frequency — plus repeat sampling if configured, which multiplies again.
A focused keyword set across a few cities on both devices at daily refresh sits at the lighter end. Large keyword sets across many locations with repeat sampling sits considerably higher. One scoping call, a free pilot on your own keywords within 48 hours, then a fixed monthly quote. Request a quote.
Send us keywords and target locations. We return full SERPs with feature detection, AI Overview citations and pixel depth within 48 hours.
Free pilot, no card, no obligation. We'll show you sampling variance on your own keyword set.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
The EU AI Act impact on web scraping & AI training data GPAI transparency, copyright reservations, prohibited practices & a compliance checklist from Actowiz.
How a B2B supplier replaced manual tender-portal checking with an automated, filtered feed of relevant government tenders from GeM and CPP/eProcure never missing a bid deadline again.
Actowiz Solutions tracks post–World Cup 2026 travel pricing — hotel ADR & airfare normalization across host cities, event-premium decay data & lessons for travel teams.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.