Bibliographic records
The core record, structured.
- Publication and application numbers
- Kind codes and document types
- Titles and abstracts
- Claims where publicly available
- Priority, filing and publication dates
With assignees normalised and families linked, not raw office records.
One company appears as forty different assignee strings across offices, and one invention appears as thirty separate filings across jurisdictions. Counting raw records tells you nothing about a portfolio. Normalisation and family linkage are the entire job.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Patent and IP data scraping is the automated collection of publicly available intellectual property records: patent applications and grants with bibliographic data, titles, abstracts and claims where published, classification codes, inventor and assignee information, priority and filing dates, legal status, citations, and trademark records.
Patent offices publish this deliberately — disclosure in exchange for protection is the bargain — so access is straightforward. The difficulty is that raw record counts are almost meaningless for the questions people actually ask.
Assignee strings are normalised to entity identities with confidence scoring, including subsidiary mapping to parent where identifiable. Filings for the same invention are linked into families with a family identifier, size and jurisdiction list, so portfolio analysis counts inventions rather than paperwork.
Legal status carries a status_as_of date, because
public status data lags actual events and a status presented as
current is misleading.
We do not provide legal opinions. Validity, infringement, freedom-to-operate and patentability are legal assessments requiring qualified counsel and, usually, licensed analytical tools. We also do not resell licensed commercial patent databases. What we provide is accurately extracted, normalised public bibliographic data that your IP counsel or analytics tools can work from.
Portfolio mapping and technology landscape work are the largest use cases.
The core record, structured.
Who owns and who invented.
One invention, one unit of analysis.
Technology filtering that works.
With honest currency.
Influence and brand protection.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 110+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
record_key / office / record_type |
string / enum | Stable record identity, issuing office and whether patent, trademark or design | Per run |
publication_number / kind_code |
string | Office publication number and kind code, so applications and grants are distinguishable | Per run |
title / abstract |
string | Title and abstract as published, in the publication language | Per run |
assignee_published / assignee_normalised /
assignee_confidence
|
string / decimal | Assignee as recorded, normalised entity and confidence in the mapping | Per run |
family_id / family_size / family_jurisdictions
|
string / int / array | Family linkage so portfolio analysis counts inventions not filings | Per run |
cpc_codes / ipc_codes / nice_classes |
array | Classification codes for technology and sector filtering | Per run |
priority_date / filing_date / publication_date
|
date | The three dates that anchor any timing analysis | Per run |
legal_status / status_as_of |
enum / date | Status as published together with the date that status reflects | Monthly |
citations_forward / citations_backward |
int | Citation counts where the office publishes them | Monthly |
inventor_names |
array | Inventor names as published in the official record | Per run |
mark_text / mark_status |
string / enum | For trademark records, the mark and its registration status | Monthly |
Legal status always ships with status_as_of. Public status data lags actual events, sometimes by months, and presenting a status as current when it reflects an older office update is the most common way patent datasets mislead.
Office data structures differ substantially. Coverage is built to your jurisdiction and technology scope.
We collect from official office publications and public registers. We do not resell licensed commercial patent databases, and we do not provide the analytical layers those products include. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| Europe & United States | EPO and USPTO publish the deepest openly accessible bibliographic data, making these the core of most engagements. |
| Japan, South Korea & China | Very high filing volumes where transliteration inconsistency makes assignee normalisation most valuable. |
| United Kingdom & Germany | Strong national office publication alongside EPO, useful for national-route filing analysis. |
| India & GCC | Growing filing volumes with improving publication accessibility, increasingly requested for emerging-market portfolio work. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
IP and R&D teams dominate, with corporate development and investors following.
Competitor portfolio assessment is defeated by assignee name variants and by filings counted instead of inventions.
Assignee-normalised, family-linked portfolio views with classification and status, so competitor portfolios are comparable.
Portfolio insight
Identifying where technology activity is concentrating requires classification-level filing trends across offices.
Filing trends by classification, assignee and jurisdiction over time, with families counted once.
Scouting hit rate
Target assessment needs a clean view of what a company actually owns, across subsidiaries and jurisdictions.
Subsidiary-to-parent mapped portfolios with family linkage, status and expiry signals for diligence support.
Diligence speed
Competitor R&D direction is visible in filings months or years before products, if the data is normalised.
New filing alerting by competitor entity and classification, with family and priority dates for timing analysis.
Signal lead time
IP quality assessment needs family breadth, citation and status data rather than raw counts.
Family-level portfolio metrics with citation counts and status, delivered as modelling-ready panels.
Diligence confidence
Innovation research needs reproducible, normalised patent data with documented methodology.
Documented extraction with assignee normalisation confidence and family definitions stated, reproducible.
Reproducibility
Four patterns, with the outcome each is judged on.
Assignee strings are normalised to entities with subsidiary mapping, and filings are linked into families, so a competitor portfolio is counted as inventions across jurisdictions rather than as raw records.
Outcome: Portfolio comparisons that reflect invention counts rather than filing strategy artefacts.
Filing volumes are analysed by classification code, assignee and jurisdiction over time using family-level counts and priority dates, showing where activity is concentrating.
Outcome: Technology direction identified from filing behaviour ahead of product announcements.
Publications are monitored against normalised competitor entities and classification scopes, with family and priority dates included so timing can be assessed.
Outcome: Competitor R&D direction surfaced as publications appear rather than through periodic review.
Target portfolios are assembled with subsidiary-to-parent mapping, family linkage, legal status with as-of dates and expiry signals, giving counsel a clean starting inventory.
Outcome: Diligence starting from a structured inventory rather than from raw office searches.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
Portfolio analysis counted filings under a single assignee string, missing subsidiary and transliterated entity variants across offices.
Assignee normalisation with subsidiary-to-parent mapping and confidence scoring, plus family linkage so inventions were counted once.
The competitor portfolio resolved to substantially more inventions than previously understood, changing the technology response.
Each diligence process began with unstructured office searches, consuming analyst time before any assessment could begin.
Normalised, family-linked portfolio inventories with classification tagging, legal status and as-of dates delivered per target.
Counsel and analysts began from a structured inventory rather than assembling one.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Assignee normalisation is continuous modelling work, and office data structures change without notice.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Assignee normalisation is the single highest-value transformation in patent data and the one most consistently underestimated. Without it, portfolio analysis is not merely imprecise — it is wrong in a systematic direction.
Normalisation combines string similarity, co-filing patterns, inventor overlap, classification profile and, where identifiable, corporate structure information, producing a normalised entity with a confidence score. Subsidiary-to-parent mapping is delivered as a separate relationship so you can analyse at either level.
Confidence matters because the failure modes are asymmetric. Over-merging two genuinely different companies inflates a portfolio and can mislead a diligence process. Under-merging fragments a portfolio and understates a competitor. We deliver the confidence score, you set the threshold, and low-confidence mappings arrive flagged rather than either applied silently or dropped.
Our rate is about 95.8% on major filers and lower on long-tail individual and small-entity filings, where the signals are weaker. We report it by filer size rather than as one figure, because that is where the difference matters.
Patent data buyers sometimes want conclusions rather than data. In IP specifically, providing them would be both unqualified and dangerous.
status_as_of precisely so nobody
treats it as live.
An FTO opinion from a data vendor has no professional standing and no insurance behind it. If a decision made on it goes wrong, you carry the consequence entirely. Qualified IP counsel exists for exactly this reason, and any vendor blurring that line is creating risk for you while sounding more useful.
What we do is make counsel's work faster and cheaper: a normalised, family-linked, classification-tagged portfolio inventory with status dates, rather than raw office search exports. That is genuinely valuable and it is honestly within our competence.
Jurisdictions, classification scope and the entity list for normalisation are agreed first, and normalisation is tuned against your known entity variants during the pilot.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect from official patent and trademark office publications and public registers. We do not resell licensed commercial patent databases, and we do not provide validity, infringement, freedom-to-operate or patentability assessments, which require qualified counsel. Legal status is always delivered with an as-of date.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What IP, R&D and corporate development teams ask during evaluation.
Because one company appears as dozens of strings across offices — legal entity variants, subsidiary names, transliterations, abbreviations and outright typos in official records. Without normalisation, a competitor portfolio is fragmented and understated.
We normalise using string similarity, co-filing patterns, inventor overlap and classification profile, delivering a normalised entity with a confidence score and separate subsidiary-to-parent mapping. Our rate is about 95.8% on major filers and lower on long-tail small-entity filings, and we report it by filer size rather than as one figure.
Because one invention filed in fifteen jurisdictions generates fifteen records. Counting records measures filing strategy rather than inventive output, and it inflates portfolios unevenly depending on how internationally a company files.
We link filings into families with a family identifier, size and jurisdiction list, so portfolio comparisons count inventions. Both views are available — family-level for portfolio comparison, record-level where jurisdiction detail matters.
No, and we will not. Validity, infringement, freedom-to-operate and patentability are legal assessments requiring qualified counsel and claim construction against specific products.
An opinion from a data vendor has no professional standing and no insurance behind it, so if a decision made on it goes wrong you carry the entire consequence. What we do is make counsel's work faster: a normalised, family-linked, classification-tagged inventory rather than raw search exports.
As current as office publication allows, which lags actual
events — sometimes by months. Every status field ships
with status_as_of stating the date the status
reflects.
This is deliberate. Presenting a status as current when it reflects an older office update is the most common way patent datasets mislead, and it matters most in exactly the situations where accuracy is critical, such as expiry and lapse assessment.
Where offices publish it publicly, yes, in the publication language. Coverage and format vary by office, and some make full text available while others publish only bibliographic data and abstracts openly.
We state per office what full-text availability looks like during scoping. Where full text is only available through a licensed product, we say so rather than substituting an abstract and letting you assume you have claims.
Yes — trademark records with mark text, Nice classes, status and opposition indications where published, plus design registrations from offices that publish them openly.
Trademark data is often more immediately useful for brand protection work, and it joins to our seller monitoring service for enforcement workflows against unauthorised sellers using registered marks.
Original script is retained alongside transliteration, and normalisation accounts for inconsistent transliteration across offices — which is a major source of assignee fragmentation for Japanese, Korean and Chinese filers.
Titles and abstracts are captured in the publication language with translation available for triage. As with tender documents, translation is for filtering rather than for analysis, and the original remains authoritative.
Different rather than cheaper, and worth being clear about. Commercial IP databases include analytical layers, curated data and search interfaces that we do not provide, and for many IP teams that product is the right purchase.
We are a better fit when you need normalised bibliographic data delivered into your own systems for integration with other datasets, or when your requirement is continuous monitoring feeding an internal pipeline rather than analyst desktop search. If your team primarily needs desktop search and analytics, licence the commercial product.
We quote individually. The drivers are jurisdiction and office count, classification or entity scope, whether full text is required where available, and refresh frequency for status and new publication monitoring.
A focused competitor entity set across major offices sits at the lighter end. Broad classification landscapes across many jurisdictions with full text and monthly status refresh sits higher. One scoping call, a free pilot on your own entity list within 48 hours, then a fixed quote. Request a quote.
Send us competitor entities or a classification scope. We return normalised, family-linked records with status and as-of dates within 48 hours.
Free pilot, no card, no obligation. Send your known entity variants and we'll tune normalisation against them.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
The EU AI Act impact on web scraping & AI training data GPAI transparency, copyright reservations, prohibited practices & a compliance checklist from Actowiz.
How a B2B supplier replaced manual tender-portal checking with an automated, filtered feed of relevant government tenders from GeM and CPP/eProcure never missing a bid deadline again.
Actowiz Solutions tracks post–World Cup 2026 travel pricing — hotel ADR & airfare normalization across host cities, event-premium decay data & lessons for travel teams.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.