How Actowiz Solutions built a daily news aggregation pipeline across 24 news and forum sites — clean article extraction, dedup, metadata and structured delivery.
A US organisation running a media-intelligence function — the kind of operation that needs to know, every morning, what was published overnight across a defined set of sources that matter to its domain. Their source list was specific and mixed: 24 websites spanning mainstream news outlets, trade and industry publications, and community forums. Their need was equally specific: a clean, daily, structured feed of everything newly published, delivered in a consistent shape they could analyse, search, and route — regardless of how differently those 24 sites were built.
The brief was, on its surface, one of the most common data requests in existence: "collect the articles from these sites, every day." Its difficulty — and the reason it's worth documenting — is that "these sites" spanned 24 completely different content architectures, and "clean" turned out to be doing enormous work in that sentence.
News and article extraction is deceptively hard, and the difficulty scales with source diversity:
Per-site extractors tuned to each of the 24 sources' structures, all feeding one normalised schema: source, URL, headline, author(s), publish datetime (normalised to a standard timezone), clean body text, content type (article vs forum thread), section/category, and collection timestamp. Twenty-four dialects in, one clean language out.
Content extraction that isolates the article or post body and strips navigation, ads, related-content rails, cookie banners, and widgets — delivering clean text suitable for reading, search indexing, or NLP, not a soup of page furniture. This is the layer that makes the dataset actually usable downstream.
For the forum sources, thread-aware extraction: original post plus structured replies, per-author and per-timestamp, with conversation structure preserved — so forum data arrives as the branching discussion it is, not flattened into a fake single article.
Daily change detection per source identifies genuinely new publications, with URL-and-content deduplication (including cross-source dedup where syndicated content appears on multiple sites) so the client's feed contains each item once, no gaps and no repeats.
Publish dates parsed from every format and timezone into consistent sortable datetimes; author attribution standardised; sections and categories mapped where the client wanted consistent taxonomy across sources.
The full day's new content delivered on schedule in structured form (the client's preferred JSON/CSV), individually-filed or as a searchable consolidated set per their workflow — with per-record lineage.
Public, published content collected responsibly with respectful request pacing per our ethical-load standards; content used for the client's monitoring and analysis with appropriate handling of source attribution; personal data in forum content (usernames, any incidental PII) masked at the edge per our compliance framework; copyright-aware handling (the pipeline collects and structures for the client's internal analysis, with attribution and source links preserved — the responsible posture for media data).
{
"record_id": "news-2026-08-11-118204",
"source": "sample_trade_pub",
"content_type": "article",
"url": "https://…",
"headline": "Sample headline text",
"authors": ["A. Writer"],
"published_at": "2026-08-11T04:30:00-04:00",
"section": "industry",
"body_text": "[clean article body, boilerplate removed]",
"word_count": 812,
"collected_at": "2026-08-11T06:05:00-04:00",
"lineage_id": "lin-2044-n"
}
| Source Type | Sites* | New Items* | Dedup Removed* | Avg Freshness Lag* |
|---|---|---|---|---|
| Mainstream news | 9 | 214 | 18 | < 2 hrs |
| Trade/industry | 10 | 96 | 4 | < 4 hrs |
| Forums | 5 | 140 threads | 6 | < 3 hrs |
Sample data — illustrative of deliverable format. Actual delivery is per-article/per-thread with full clean text and normalised metadata.
| Metric | Value* |
|---|---|
| Sources | 24 (news, trade, forums) |
| Content types | Articles + forum threads (native handling) |
| New items per day (typical) | 400–500 |
| Body-extraction cleanliness (audited) | 98%+ boilerplate-free |
| Delivery | Daily, scheduled, client schema |
| Missed-publication rate (audited) | Under 1% |
| Site changes absorbed, first quarter | 14 (13 auto-repaired) |
| Time to full 24-source pipeline live | 3 weeks |
Representative engagement figures — illustrative of project structure.
The client got the thing a media-intelligence function actually runs on: a complete, clean, consistent daily feed that let their analysts start each morning with the day's relevant content already collected, deduplicated, dated, and structured — rather than 24 browser tabs and a copy-paste routine. The boilerplate-free body text meant the feed plugged directly into their search and analysis tools without a cleaning step; the native forum handling meant community discussion (often the earliest signal in their domain) arrived as usable structured conversation; and the sub-1% missed-publication rate meant they could trust the feed as complete rather than treating it as a starting point to double-check.
The consolidation of 24 disparate sources into one schema was the quiet transformation: cross-source analysis (what's being said across outlets and forums about a topic, today) became a query rather than a manual assembly job. And the self-healing layer meant the morning feed kept arriving complete even as sources redesigned — the reliability that turns a data feed from a tool into infrastructure.
The engagement continues with sources added over time onto the same schema, and an enrichment layer (topic tagging and sentiment) explored as a next phase — the pipeline designed so new sources and new processing slot in without disrupting the daily delivery the client now depends on.
Media monitoring, competitive intelligence, research, and content aggregation all rest on the same need: a defined set of sources, collected completely and cleanly, every day, into one consistent shape. The transferable design: source-aware extractors feeding a unified schema, boilerplate-free body extraction, native handling of different content types (article vs thread), reliable new-content detection with cross-source dedup, metadata normalisation, and self-healing collection so a daily-depended-upon feed never quietly arrives incomplete. The value is in completeness and cleanliness — the two things manual and naive approaches fail at first.
Yes — content extraction isolates the article or post body and strips boilerplate (navigation, ads, related rails, cookie notices, widgets), delivering clean text suitable for search indexing and NLP.
Natively — forum threads are extracted as original-post-plus-structured-replies with per-author, per-timestamp, conversation-structure preserved, rather than flattened into a single article body.
Daily new-content detection per source with URL-and-content deduplication (including cross-source dedup for syndicated content) ensures each item appears once, with an audited sub-1% miss rate.
Yes — new sources slot into the same unified schema without disrupting delivery. Contact Actowiz Solutions to scope a daily media pipeline for your source set.
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Unlock retail insights with Noon Saudi Arabia Product Data Extraction to track prices, inventory, discounts, and product trends in real time.
Unlock Hertz & Avis Rental Car Data for Dynamic Pricing Intelligence to track rental rates, availability, and market trends in real time.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.