Emissions disclosures
Reported figures with the qualifiers that matter.
- Scope 1, 2 and 3 by category
- Units and reporting standard
- Consolidation boundary
- Intensity metrics and denominators
- Restatement tracking with prior values
Reported figures with page references, not a score we invented.
An ESG score is somebody's opinion expressed as a number, and the methodology is usually proprietary. What your analysis actually needs is the reported figure, the page it came from, and whether anyone assured it. That is what we deliver.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
ESG data scraping is the automated extraction of sustainability information that companies publish: greenhouse gas emissions by scope, energy and water use, waste figures, workforce and diversity data, governance disclosures, targets and baselines, assurance statements, and product-level sustainability claims.
The disclosures are public, usually in PDF reports, and almost entirely unstructured. Extracting them accurately with traceability is the work. Turning them into a score is not something we will do.
Commercial ESG ratings exist and are licensed products. If you need them, licence them. We provide the disclosed figures, traceable to source, so your own methodology can be applied and defended.
All three are delivered as fields, because a tonnage figure without them cannot support any conclusion.
Emissions extraction is the largest use case. Claims capture is the fastest growing, driven by greenwashing regulation.
Reported figures with the qualifiers that matter.
What is promised, precisely.
The evidential quality layer.
Reported people metrics.
Structure as published.
Where greenwashing risk sits.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 130+ metrics and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
company / company_id |
string | Company as published and normalised identity for longitudinal joins | Per report |
metric / period |
enum / string | Normalised metric identifier and the reporting period it covers | Per report |
value / unit |
decimal / string | Reported figure and its unit exactly as published, not converted silently | Per report |
reporting_standard / boundary |
enum | Standard followed and consolidation boundary, both of which change the total | Per report |
restated / restated_from |
boolean / decimal | Whether the figure restates a prior disclosure, and the prior value | Per report |
assurance |
object | Presence, level, whether the provider is named and which metrics were in scope | Per report |
target_year / baseline_year / scopes_covered
|
int / array | Target structure including which scopes are and are not covered | Per report |
third_party_validated |
boolean | Whether a target is externally validated, where the company states it | Per report |
source_doc / source_page / source_table |
string / int | Exact provenance for every figure, so any number is auditable | Per report |
extraction_confidence |
decimal | Confidence in the extraction, with low-confidence figures flagged not dropped | Per report |
claim_text / claim_first_seen |
string / date | Sustainability claims as worded, with first-observed date for change tracking | Weekly |
Units are retained exactly as published rather than silently converted. Where conversion is needed we supply factors separately, because a silent unit conversion is the most common way emissions comparisons break.
Disclosure requirements differ by jurisdiction and are changing. Coverage is built to your company universe.
We extract what companies publish. We do not verify whether disclosed figures are accurate, and we do not assess whether a claim is substantiated — both are assurance functions, not extraction functions. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| European Union & United Kingdom | CSRD-aligned reporting created the deepest structured disclosure base, which makes extraction most complete here. |
| United States | Climate disclosure in filings plus voluntary sustainability reporting, with substantial variation in boundary and assurance practice. |
| Japan & Australia | Strong ISSB and TCFD-aligned reporting adoption with good assurance disclosure. |
| India & GCC | Rapidly expanding mandatory sustainability reporting with growing disclosure depth year on year. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Sustainability and investor teams dominate, with supply chain and regulatory analysts growing.
Peer benchmarking requires comparable disclosed figures, and reports publish them inconsistently across boundaries and standards.
Peer emissions and target data with boundary, standard, restatement and assurance captured, traceable to source pages.
Peer benchmark quality
Ratings disagree and hide their inputs, so you need the disclosed figures to apply your own methodology.
Reported figures with full qualifiers and page-level traceability across your investment universe, with no score applied.
Methodology defensibility
Supplier emissions and target claims need collecting from hundreds of supplier reports, manually today.
Supplier-level emissions, targets and assurance status extracted from their published reports, refreshed as reports publish.
Scope 3 data coverage
Assessing claim substantiation across a market requires collecting claims as worded, with dates.
Product and corporate sustainability claims captured verbatim with first-seen dates and change detection.
Market surveillance coverage
You need to know how peers disclose in order to decide your own disclosure boundaries and detail.
Peer disclosure practice comparison — which metrics, which boundaries, what assurance level, what is omitted.
Disclosure completeness
ESG research needs reproducible, source-traceable disclosure data rather than vendor scores.
Documented extraction with page-level provenance and stated methodology, reproducible and auditable.
Reproducibility
Four patterns, with the outcome each is judged on.
Reported emissions are extracted with unit, standard, boundary, restatement status and assurance level, so peer comparison holds rather than comparing figures computed on different bases.
Outcome: Benchmarking that survives scrutiny because every figure carries its qualifiers and page reference.
Target years, baselines, scopes covered and exclusions are captured, revealing which commitments exclude Scope 3 and which are externally validated.
Outcome: Target credibility assessed on structure rather than on headline year.
Supplier published reports are processed for emissions, targets and assurance status, building the Scope 3 supplier picture that manual collection cannot cover.
Outcome: Supplier ESG coverage expanded without proportional analyst headcount.
Product and corporate sustainability claims are captured verbatim with first-seen dates and change detection, so claim evolution and quiet withdrawals are visible.
Outcome: Claim inventory with dated wording, which is what a substantiation review requires.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Two providers gave materially different scores for the same holdings, and neither exposed the disclosed figures behind them, blocking the firm's own methodology.
Extraction of reported emissions, targets and assurance status with units, boundaries and page-level references, with no score applied.
The firm applied its own methodology to disclosed figures and could evidence every input to its investment committee.
Sustainability reporting benchmarked peer emissions without accounting for consolidation boundary or restatement, producing comparisons that did not hold.
Extraction with boundary, reporting standard, restatement flags and prior values retained as separate records, plus assurance level.
Benchmarking moved onto comparable bases, and several apparent peer advantages proved to be boundary differences.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
PDF table extraction with page-level traceability across hundreds of reports is specialist, repetitive work.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Emissions restatement is routine, frequently substantial, and the single most common cause of broken ESG trend analysis. It also receives almost no attention in most datasets.
A dataset that overwrites prior values with restated ones shows a smooth trend that never existed. A dataset that keeps only originally reported values compares figures computed on different bases. Both produce confident, wrong reduction percentages.
Every figure carries restated and, where
identifiable, restated_from with the prior value. We
keep both the original and the restated disclosure as separate
records with their own source references, so you can construct
either an as-reported series or a restated series — and see
where they diverge.
That divergence is itself informative. A company restating substantially in the same year it announces a reduction is worth a closer look, and only restatement tracking makes that visible.
Sustainability disclosures live in PDF reports running to hundreds of pages, with figures in tables, footnotes and narrative text. Extraction is the technical core of this service, and traceability is what makes the output usable.
ESG figures get challenged — by investors, auditors, regulators and journalists. A tonnage number without provenance is an assertion. The same number with a document name, page and table caption is evidence someone can verify in thirty seconds.
Our extraction confidence runs at about 93%, and low-confidence figures are flagged rather than dropped, so an analyst can check the specific page rather than discovering a gap. Units are retained exactly as published, with conversion factors supplied separately, because silent unit conversion is the most common way emissions comparisons break.
For monitoring controversies and reported incidents alongside disclosures, this pairs with our news data service, which uses the same company entity resolution.
Company universe and metric scope are agreed first, since report count and metric breadth drive extraction effort.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We extract from publicly published corporate reports, regulatory filings and company websites. We do not verify whether disclosed figures are accurate, assess whether claims are substantiated, or produce ESG ratings, scores or rankings. Every extracted figure carries a document and page reference for independent verification.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What sustainability, investment and compliance teams ask during evaluation.
No, deliberately. A score requires weighting incommensurable things — emissions against board diversity against water use — and every weighting is an opinion rather than a measurement. The same company routinely gets very different scores from different providers, which tells you the score measures the methodology as much as the company.
We extract the disclosed figures with their qualifiers and page references so you can apply and defend your own methodology. Commercial ratings exist as licensed products; if you need them, licence them.
Because overwriting prior values with restated ones produces a smooth trend that never existed, and keeping only originals compares figures computed on different bases. Both produce confident, wrong reduction percentages.
We keep both as separate records with their own source
references and a restated_from value where
identifiable, so you can build an as-reported or restated
series and see where they diverge. That divergence is often
the interesting finding.
About 93%, with low-confidence figures flagged rather than dropped so an analyst can check the specific page. Every figure carries document, page and table caption references.
Extraction is genuinely hard here: multi-level table headers, footnote markers inside values, figures split across pages, inconsistent units within a single report, and qualifiers disclosed only in small print beneath tables. We report the rate rather than a rounded claim because the residual is where the analyst effort goes.
No. Verification is an assurance function requiring access to underlying company data, which we do not have and would not claim to. We extract what is published, exactly as published.
What we do capture is whether a figure was assured, at what level, whether the provider is named, and which metrics were in assurance scope. That is the evidential quality signal available from public disclosure, and it is more useful than an unverifiable accuracy claim from us.
We capture claims as worded with first-seen dates and change detection, including quiet withdrawals and rewordings. We do not assess whether a claim is substantiated.
Substantiation requires evidence about the underlying activity, which is not in the claim. What we provide is a dated claim inventory — exactly what a substantiation review or regulatory assessment starts from. Concluding on substantiation is your assessment or a regulator's, not our pipeline's.
By capturing which of the fifteen categories a company reports, with their labels as published, rather than forcing them into a single total. Companies report different subsets with different labels, and summing incomparable subsets produces a meaningless figure.
Where a company reports a Scope 3 total without category breakdown, we record that and flag the absence. A total covering four categories is not comparable to one covering eleven, and the field structure makes that visible.
Yes, where suppliers publish. Coverage is thinner because smaller and private companies disclose less, and some publish nothing at all.
We give you a realistic coverage estimate for your supplier list during scoping rather than promising universal coverage. For suppliers that publish nothing, no extraction service can help — that gap needs a supplier engagement programme, not a data vendor.
Report-driven rather than on a fixed cadence, since sustainability reports publish annually with some interim disclosures. We monitor for new report publication and process on release.
Claims monitoring on websites and product pages runs weekly, because claim wording changes without any report cycle and quiet withdrawals happen between reports.
We quote individually. The drivers are company universe size, metric breadth, whether claims monitoring is included, and how much of the extraction requires PDF table work versus structured filings.
A focused peer group with core emissions and target metrics sits at the lighter end. Large universes with broad metric coverage and continuous claims monitoring sits higher. One scoping call, a free pilot on your own company list within 48 hours, then a fixed quote. Request a quote.
Send us companies and metrics. We return extracted figures with units, boundaries, assurance status and page references within 48 hours.
Free pilot, no card, no obligation. No scores, no ratings — disclosures with provenance.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
The EU AI Act impact on web scraping & AI training data GPAI transparency, copyright reservations, prohibited practices & a compliance checklist from Actowiz.
How a B2B supplier replaced manual tender-portal checking with an automated, filtered feed of relevant government tenders from GeM and CPP/eProcure never missing a bid deadline again.
Actowiz Solutions tracks post–World Cup 2026 travel pricing — hotel ADR & airfare normalization across host cities, event-premium decay data & lessons for travel teams.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.