Creative & copy
The ad itself.
- Headline and body copy as published
- Format: image, video, carousel, text
- Variant counts within an ad set where disclosed
- Call-to-action label
- Creative reference for archival and change detection
From public ad libraries, with spend estimates deliberately excluded.
Platform ad libraries are the most under-used public dataset in marketing. Every active ad your competitor is running is disclosed, with creative and run dates. What is not disclosed is what they paid, and no amount of scraping produces that.
Free pilot on your own sources, returned in 24 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Ad intelligence data scraping is the automated collection of what advertising platforms publicly disclose about active advertising: the creative itself, ad copy, formats, the countries an ad was shown in, when it started and last appeared, the advertiser, and where it links to.
These libraries exist because of transparency regulation and platform policy, and they are genuinely public. That makes this one of the more comfortable collection categories we run — and one of the most under-used, because most teams do not realise how much is disclosed.
Vendors do sell commercial spend estimates. Those are models — derived from panels, sampled exposure and proprietary assumptions — not observations. Presenting a model inside a dataset labelled as ad library data is the same problem as app revenue estimates or ESG scores, and we take the same position.
spend_estimate ships as a constant not_collected, and impressions as not_published_commercial, so the exclusions are visible in the data rather than looking like gaps.
The observable proxy for performance is how long an ad runs. Advertisers pause creative that does not work. An ad running fifty-two days with six variants is being invested in; one that ran four days and stopped was probably a test that failed. That is not spend, but it is evidence, and it is fully observable.
Creative longevity and variant counts are where the analytical value sits. Landing page capture is the most under-used.
The ad itself.
The observable performance proxy.
Who is running it.
Only what is published.
Where the money is pointed.
Where more is published.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 80+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
ad_id / library |
string / enum | Record identity and which public library it came from | Every run |
advertiser_name / advertiser_verified |
string / boolean | Advertiser as disclosed and verification status where shown | Every run |
country_shown |
array | Countries the ad was shown in, where the library discloses it | Every run |
first_seen / last_seen / days_running |
date / int | Run window and computed duration, the observable performance proxy | Daily |
still_active |
boolean | Whether the ad was live at last observation | Daily |
format / variants |
enum / int | Creative format and variant count where disclosed | Every run |
headline / cta_label |
string | Ad copy and call-to-action label as published | Every run |
landing_domain / landing_path |
string | Campaign destination, so what is being promoted is measurable | Every run |
creative_archived |
boolean | Whether the creative reference was captured for change detection | Every run |
spend_estimate / impressions |
constant | not_collected and not_published_commercial, stated in the data | Every run |
spend_range_published / impressions_range_published |
string | For political and issue ads only, the ranges the platform publishes | Where applicable |
spend_estimate is a constant not_collected. Vendors do sell commercial spend estimates; they are models built on panels and assumptions rather than observations, and we will not blend one into a dataset labelled as ad library data.
Coverage follows what platforms publish. Disclosure depth varies considerably between libraries and by ad category.
Disclosure depth differs sharply between libraries and by ad category. We state per library what is actually published before build rather than implying uniform coverage. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United States | Deepest ad library disclosure and the largest advertiser base, which makes competitive creative analysis most complete here. |
| European Union & United Kingdom | Transparency regulation drives the fullest disclosure, including regulated-category spend and impression ranges. |
| India | Very high advertiser volume with rapid creative turnover, where longevity analysis separates tests from proven creative. |
| Southeast Asia & GCC | Growing disclosure coverage as libraries extend to more markets, with strong short-form creative activity. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Brand marketing and agency strategy teams dominate, with competitive intelligence close behind.
You cannot see what competitors are running, for how long, or what they are promoting, despite it being publicly disclosed.
Active and historical creative with run dates, variant counts and landing destinations across your competitive set.
Creative test velocity
Creative direction is set from intuition and a handful of screenshots rather than from the full disclosed set.
Full creative and copy capture with longevity data, so what competitors keep running is separable from what they tested and dropped.
Creative win rate
Competitor campaign launches are noticed late, and promotional messaging shifts are missed entirely.
Daily new-ad detection by competitor with copy capture and landing page classification.
Response time
You cannot tell which products or collections competitors are actively pushing.
Landing domain and path capture with destination classification, revealing which SKUs and collections are being promoted.
Assortment response
Category messaging trends need systematic copy analysis rather than anecdote.
Ad copy corpus by category and market with longevity weighting, so persistent messaging is distinguishable from experiments.
Insight coverage
Political and issue advertising research needs reproducible archive collection with published ranges retained.
Regulated-category ads with published spend and impression ranges, funding entities and disclaimer text.
Reproducibility
Four patterns, with the outcome each is judged on.
Run dates and active status produce days-running per ad, so creative a competitor keeps investing in is separable from tests they dropped within a week.
Outcome: Creative direction informed by what competitors sustain rather than by what they launched.
New ads are detected daily per advertiser with copy captured, so campaign launches and messaging shifts surface as they happen.
Outcome: Competitive response measured in days rather than in quarterly reviews.
Landing domain and path capture reveals which products, collections or offers competitors are actively driving traffic to.
Outcome: Assortment and promotional response based on where competitors point spend, not where they say they focus.
Political and issue ads are collected with the spend and impression ranges platforms publish, plus funding entities and disclaimer text.
Outcome: Reproducible archive research with published figures retained rather than modelled.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
The brand team collected competitor ads manually and had no view of which creative was sustained versus tested and dropped.
Ad library collection with run dates, variant counts and computed days-running across the competitor set.
Sustained creative became separable from failed tests, changing how creative direction was set.
Available spend figures were vendor models the agency could not explain or stand behind in a client meeting.
Creative longevity and variant counts as the observable performance proxy, with spend explicitly excluded and the reason documented.
Client reporting moved to observable evidence the agency could defend rather than modelled figures it could not.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Libraries change structure and disclosure rules frequently, and creative archival is storage-heavy.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 24 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Everyone buying ad intelligence wants to know what is working. Spend is not published, performance is not published, and the estimates on the market are models. But there is a real, fully observable signal, and it is under-used.
Longevity is a proxy, not a measurement. A brand may run creative for brand reasons rather than performance ones, and library dates can be imprecise where a platform reports coarsely. We deliver first_seen, last_seen and days_running from library-published dates, and where a library reports a date range rather than a date we retain the range rather than picking a midpoint.
It is still the strongest observable signal in this category, and it costs nothing beyond collecting the dates properly.
Ad library disclosure is uneven in a way that matters for scoping, and it is worth understanding before you commit to a comparison across categories.
Creative, copy, format, run dates, advertiser and often geography. No spend, no impressions, no reach, no targeting. This is the majority of what we collect and it is genuinely rich — it just does not include money.
Regulation requires more. Platforms publish spend ranges, impression ranges, funding entities and disclaimer text. We capture the published ranges as ranges — not midpoints, not point estimates. A range of five to ten thousand is what was disclosed, and converting it to seven and a half thousand invents precision.
Some categories carry additional disclosure requirements depending on jurisdiction. Coverage varies and we state it per library and market rather than generalising.
If your question is about competitor creative, messaging and what they are promoting, ad libraries answer it well. If your question is share of voice in monetary terms, they do not, and no vendor's model will give you a figure you can defend in a board paper. In that case licence a measurement product and be clear it is modelled.
For paid placement inside retailers rather than on social and search platforms, that sits in our retail search and share of shelf service, where sponsored slots are detected on the retailer's own results pages.
Advertiser set, libraries, markets and whether creative archival is required are scoped first, since creative storage drives cost.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect from public ad transparency libraries and ads centres without accounts or credentials. We do not collect targeting parameters, audience definitions or performance data, none of which is published. We do not provide modelled spend or impression estimates for commercial advertising. Published ranges for regulated categories are retained as ranges.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 24 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What marketing, agency and research teams ask during evaluation.
No for commercial advertising, because platforms do not publish it. spend_estimate ships as a constant not_collected so the exclusion is visible in the data.
Vendors do sell spend estimates — those are models built on panels and proprietary assumptions, not observations. If you need a spend figure, licence a measurement product and label it as modelled. We will not blend a model into a dataset labelled as ad library data.
Creative longevity, which is fully observable. Advertisers pause creative that fails, so days-running is a revealed preference. An ad live for fifty-two days with six variants is being invested in; one that ran four days was probably a failed test.
The distribution matters most: an advertiser whose median ad runs four days is testing heavily, one whose median runs forty is running proven creative. Those are different marketing operations, and the comparison is fair because the dates come from the same disclosure for everyone.
Yes, for political and issue advertising, where regulation requires platforms to publish spend and impression ranges. We capture those as ranges.
We do not convert a five-to-ten-thousand range into a midpoint. That invents precision the disclosure does not contain, and for research use in particular the range is the finding.
We capture copy, headline, format, variant count and CTA label as text, plus a creative reference for change detection, at about 97% coverage.
Where creative archival is required — storing the image or video itself — that is scoped separately because storage drives cost materially. Most clients find copy and metadata sufficient for analysis and archive selectively.
No. Targeting parameters and audience definitions are not published in ad libraries. We capture the countries an ad was shown in where a library discloses that, and nothing beyond it.
We do not infer audience from creative, which is a thing some tools do. An inferred audience presented as data is a guess wearing a field name.
Major social platform ad libraries, search platform ads transparency centres, video platform disclosures and short-form creative centres where public.
Disclosure depth differs sharply between them, so we state per library what is actually published before build. A comparison across libraries needs to account for those differences rather than assuming parity.
These libraries exist specifically for public transparency, which makes this one of the more comfortable categories we operate in. Some platforms also provide official APIs for their libraries, which we use in preference to page collection where available.
We respect stated rate limits, collect without accounts, and provide a written methodology document per library plus a DPA before signature.
Yes — landing domain, path, any campaign parameters present in the URL, and destination page type. This is the most under-used field in ad intelligence.
It answers a question creative alone cannot: which products, collections or offers a competitor is actually pointing spend at. Landing page change detection on a running ad often signals a promotional shift before the creative changes.
We quote individually. Drivers are advertiser count, library count, market count, refresh frequency and whether creative archival is required — archival is the main cost variable because of storage.
A focused competitor set across two libraries in one market at daily refresh sits at the lighter end. Broad advertiser coverage across libraries and markets with full creative archival sits considerably higher. One scoping call, a free pilot on your own competitor set within 24 hours, then a fixed monthly quote. Request a quote.
Send us a competitor list. We return their active and recent ads with copy, run dates, variant counts and landing pages within 24 hours.
Free pilot, no card, no obligation. No modelled spend figures — those are not published by anyone.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Price intelligence fails at product matching, not at collection. A practical guide to the five layers of a working programme, what to measure, and how to scope a first phase.
One product category, named competitor brands, several countries, weekly refresh, delivered as data plus a Power BI dashboard. How narrow-and-deep beats broad-and-shallow.
Fliggy hotel and flight price monitoring helps travel businesses track fares, hotel rates, availability, and competitor pricing for smarter decisions.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.