How Actowiz Solutions built an automated, keyword-driven news intelligence pipeline monitoring 24 leading global publishers daily — with historical backfill, incremental delta delivery, and multilingual article capture — delivered as a standardised UTF-8 CSV dataset.
A global news and media-intelligence client needing continuous, keyword-driven monitoring of 24 premium and public publishers, delivered daily as standardised CSV.
Client name withheld under NDA.
The client operates in the news and media-intelligence space and required continuous, structured visibility into how specific topics are covered across leading international publishers. The goal was a scalable, keyword-driven monitoring capability spanning finance, energy, construction, politics, and other verticals — with clean, analysis-ready article data delivered every day.
Actowiz Solutions designed and deployed an automated daily extraction framework targeting 24 premium and public news websites across multiple global regions and industries. For the initial delivery the system captured all historical articles from 1 June 2026 onward; from the second run, the pipeline operates in delta mode — extracting only newly published articles since the last successful run — ensuring incremental daily delivery without duplication.
Actowiz engineered a keyword-driven extraction framework with source-specific parsers behind a single normalisation and delivery layer, executing a repeatable 12-step process on every daily run.
Each of the 24 sources received a dedicated parser tuned to its site architecture and rendering logic. Subscription platforms were accessed through managed authenticated sessions using securely provided credentials, while public sources were handled without login — all feeding one common downstream schema.
The first execution performed a full backfill from 1 June 2026 to the current date. From the second execution onward, the pipeline compares against prior runs and extracts only newly published articles, guaranteeing daily incremental delivery with zero duplication across cycles.
For every matched article the pipeline captured the full body, headline, author, and publication metadata, plus any embedded image, video, and audio URLs. Publication names were normalised, ISO country codes applied, and headers standardised so records from all sources align to one schema.
All output is serialised as UTF-8 with BOM, comma-delimited CSV conforming to the approved 20-field schema. Unavailable attributes are left blank rather than populated with placeholders, preserving schema integrity for direct downstream ingestion.
The framework covers 24 approved global news sources across eight industry verticals — including finance, energy, construction, and politics — and multiple geographic regions across Asia, Europe, the Middle East, and the Americas.
| Caixin Global | Sankei | Nikkei |
|---|---|---|
| Business Times | Economic Times | El Mercurio |
| Argaam | Axios | Bloomberg |
| Chosun | Construction Week Online | Foreign Policy |
| SCMP | Construction Week Saudi | MEED |
| MEES | The Atlantic | Energy Intelligence |
| Straits Times | Upstream Online | FD.nl |
| Volkskrant | Petroleum Economist | The Times |
The pipeline executes the following 12-step process at each daily run:
| Step | Process | Description |
|---|---|---|
| 1 | Input Processing | Read the keyword sheet; detect and ignore blank keywords |
| 2 | Website Selection | Iterate through all 24 approved news sources |
| 3 | Authentication | Access subscription sources using provided credentials |
| 4 | Keyword Search | Query each source with valid keywords to identify matching articles |
| 5 | Article Collection | Capture all matching article URLs for the current cycle |
| 6 | Article Extraction | Extract full body, headline, author, and publication metadata |
| 7 | Media Extraction | Identify and capture image, video, and audio URLs where available |
| 8 | Metadata Standardisation | Normalise publication names, apply ISO country codes, standardise headers |
| 9 | Historical / Delta Logic | Full backfill on first run; delta extraction on all subsequent runs |
| 10 | QA Validation | Validate mandatory fields, check for blanks, verify URL integrity |
| 11 | CSV Generation | Produce UTF-8 (BOM), comma-delimited CSV conforming to the schema |
| 12 | Delivery | Share the validated dataset via the agreed daily mechanism |
Each article record in the daily CSV contains the following 20 standardised fields:
| # | Field Name | Description |
|---|---|---|
| 1 | Title | Article headline as published on the source website |
| 2 | Message | Clean, complete article body text — free of HTML and formatting artefacts |
| 3 | Created Time | Article publication date and time |
| 4 | Source | Fixed value: News |
| 5 | Domain | Root domain of the source (e.g., nikkei.com) |
| 6 | Publication Name | Name of the publishing outlet |
| 7 | Journalist | Article author or journalist name, where available |
| 8 | Country | ISO country code for the source or article origin |
| 9 | Media Title | Media or page title associated with the article |
| 10 | Conversation Stream | Full article body text — same content as the Message field |
| 11 | Permalink | Permanent canonical URL of the article |
| 12 | Reach | Estimated audience reach of the source, where available |
| 13 | Web Shares | Social or web share count, where available |
| 14 | Engagement (Other) | Views, comments, reactions, or other metrics, where available |
| 15 | Keyword | The client-provided keyword that matched this article |
| 16 | Article ID | Unique article identifier from the source platform, where available |
| 17 | Image URLs | All image URLs associated with the article |
| 18 | Video URLs | All video URLs embedded in or associated with the article |
| 19 | Audio URLs | All audio URLs embedded in or associated with the article |
| 20 | Additional Fields | Any other available structured attributes not covered above |
Every daily dataset was validated against a multi-layer framework before delivery:
| Validation Check | Rule Applied |
|---|---|
| Mandatory field completeness | Title, Message, Created Time, Permalink, and Keyword are always populated |
| Blank-field handling | Unavailable attributes left blank — never filled with placeholder values |
| URL integrity | Permalinks and media URLs verified as well-formed and accessible |
| Delta de-duplication | No article re-delivered across runs; only new items since the last cycle |
| Country code validation | Country populated with a valid ISO code per source or article origin |
| Publication normalisation | Publication names standardised to a consistent canonical form |
| Body-text cleaning | Message and Conversation Stream stripped of HTML and junk characters |
| Encoding compliance | Output validated as UTF-8 with BOM, comma-delimited per the schema |
| Schema conformance | Final CSV validated against the approved fixed 20-field header set |
| Metric | Value |
|---|---|
| Industry | News & Media Intelligence |
| Coverage | Global — Asia, Europe, Middle East, Americas |
| Target Sources | 24 approved news websites (premium and public) |
| Input | Client-provided keyword sheet |
| Historical Coverage | From 1 June 2026 |
| Delivery Mode | Daily offline CSV (Phase 1); live feed planned for Phase 2 |
| Output Format | CSV — UTF-8 with BOM, comma-delimited |
| Output Schema | Fixed 20 fields |
| Frequency | Daily — once every 24 hours |
| Setup Timeline | 7 to 8 working days |
Actowiz Solutions designs custom, large-scale scraping, extraction, and API-delivery pipelines with rigorous QA. Visit actowizsolutions.com to discuss your data requirement.
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Zeptos 2026 IPO explained through data market share, dark stores, revenue and how Blinkit, Zepto and Instamart compare. See the numbers.
Track pricing, inventory, promotions, assortment, and dark-store availability across UAE and GCC quick-commerce platforms. Gain real-time intelligence to optimize retail strategies and outperform competitors.
KSA quick commerce mapped from public data — Nana, Rabbit, Jahez & HungerStation coverage zones, pricing & assortment across Saudi cities.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.