Posting core
The listing itself, normalised.
- Raw and normalised job title
- Seniority and function classification
- Employment type and contract length
- Work mode: onsite, hybrid, remote
- Posting and closing dates
Deduplicated across boards, because the same role is posted eleven times.
One vacancy appears on the company career page, three aggregators, two niche boards and five reposts. Count them raw and you conclude a company is hiring eleven people. Deduplication is not a nice-to-have here — it is the entire difference between signal and noise.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Job data scraping is the automated collection of publicly posted vacancy information: job title, employer, location, work mode, employment type, salary where published, required skills, seniority and posting dates. It is one of the most widely used alternative datasets, because hiring is a leading indicator of almost everything a company is about to do.
It is also the category where raw collection is most misleading, and the reason is structural: job postings are deliberately duplicated.
Our observed ratio is around 3.9 raw postings per genuine vacancy, and it varies enormously by sector — logistics and retail run far higher. Any analysis built on raw counts overstates hiring by roughly a factor of four, unevenly.
Postings are clustered on employer identity, normalised title, location and description similarity, then collapsed to one canonical record. We keep sources_seen as a field, because posting breadth is itself a signal — a role pushed to eight boards is usually harder to fill than one sitting only on a career page.
Reposts are linked rather than merged: repost_of points at the earlier posting, so you can measure genuine time-to-fill instead of a reset clock. And we distinguish agency-posted roles from direct employer postings, since agency listings frequently obscure the actual employer.
Most engagements start with a company watchlist or a role taxonomy, then extend into salary and skills analysis.
The listing itself, normalised.
Who is actually hiring, including behind agencies.
Where published, parsed properly.
The demand-side view of capability.
Where the work actually is.
The signal layer investors and planners use.
A managed engagement, not a tool licence. We own the pipeline and everything that breaks in it.
Every engagement delivers a documented schema. These are the core fields; the full dictionary runs to 110+ and is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
posting_key |
string | Canonical posting identity after cross-board deduplication | Every run |
title_raw / title_normalised |
string | Title as posted and mapped to a standard taxonomy for comparison | Every run |
seniority / function |
enum / string | Seniority level and functional area separated from the raw title | Every run |
company / company_domain |
string | Employer as posted and resolved domain, which enables joining to firmographics | Every run |
is_agency_posting |
boolean | Whether the listing came from a staffing agency rather than the employer directly | Every run |
location / work_mode |
string / enum | Location as posted with geocoding, plus onsite, hybrid or remote classification | Every run |
salary_min / max / currency / period |
decimal / enum | Parsed salary range with currency and period, null where not published | Every run |
skills |
array | Required and preferred skills extracted from the description, flagged separately | Every run |
first_seen / closed_at / days_open |
date / int | Lifecycle dates, with days open measured from first observation | Daily |
sources_seen / canonical_source |
int / enum | How many boards carried the role, and which source we treated as canonical | Daily |
repost_of |
string | Link to an earlier posting for the same vacancy, so time-to-fill is not reset | Daily |
Salary is published on roughly 61% of postings in our coverage, and that share varies sharply by market and by whether pay transparency law applies. We report the fill rate rather than imputing a range, because an imputed salary in a compensation benchmark is worse than a gap.
Coverage is built around your watchlist or taxonomy. Company career pages are usually the highest-value source and the most work to maintain.
Applicant tracking system boards — Greenhouse, Lever, Workday, Ashby — are often the best source because they are the employer's own posting, not a syndicated copy. We prefer them as canonical where available. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United States | The densest board ecosystem and widest ATS-hosted career page adoption, which makes career-page-first collection unusually effective. |
| United Kingdom & Ireland | High board density plus strong pay transparency norms, so salary fill rates are among the best available. |
| Germany, Netherlands & Nordics | Structured labour markets with good posting discipline, heavily used for skills demand and wage analysis. |
| India & GCC | Enormous posting volume with very high duplication, which is exactly where deduplication earns its cost. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
HR tech and investment analysts dominate, with workforce planning and sales intelligence close behind.
Your product needs comprehensive current listings, and maintaining board scrapers plus deduplication logic is consuming engineering capacity.
A deduplicated, normalised posting feed with stable keys and lifecycle fields into your systems, with board changes fixed by us.
Engineering hours reclaimed
Hiring is a leading indicator, but raw posting counts overstate hiring by roughly fourfold and unevenly by sector.
Deduplicated hiring velocity by company, function and geography, with repost linking so trend lines reflect real vacancies.
Signal lead time
Compensation benchmarking and skills demand analysis rely on survey data that is out of date on arrival.
Live salary ranges and skills demand by role, level and market, with the salary publication rate reported so gaps are visible.
Offer acceptance rate
Identifying which employers are hiring which roles, at what pay, requires monitoring hundreds of sources manually.
Employer-level hiring activity with agency versus direct classification, so you can see which roles are already agency-served.
Placements per consultant
Official labour statistics lag by months, and vacancy surveys sample rather than observe.
Longitudinal vacancy panels by occupation, region and skill, deduplicated so counts are comparable over time.
Analysis lead time
Hiring signals reveal budget and initiative before any announcement, but the signal is buried in duplicate noise.
Trigger events from hiring activity — new function entry, team scaling, technology mentions in requirements — on your target accounts.
Pipeline from triggers
Four patterns, with the outcome each is judged on.
Postings are deduplicated and attributed to resolved employer domains, so hiring volume by function and geography is measurable over time. Repost linking prevents refreshed listings inflating apparent growth, and closed detection gives a days-open distribution.
Outcome: Company-level hiring trends that reflect real vacancies rather than posting behaviour.
Salary ranges are parsed with currency and period and normalised to a common basis, benchmarked by normalised title, seniority and market, with the publication rate reported alongside every aggregate.
Outcome: Pay benchmarks current to the week, with the coverage gap stated rather than hidden by imputation.
Skills, tools and certifications are extracted from description text, showing which technologies are gaining requirement share by function and region — often the earliest observable signal of platform adoption.
Outcome: Technology adoption curves observed from hiring requirements ahead of vendor disclosure.
New function entry, team scaling and named technologies in requirements are surfaced as events on target accounts, filtered to exclude agency reposts and refreshed listings.
Outcome: Outbound timed to observed budget and initiative signals rather than to a cadence.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
The platform ingested postings from several aggregators without deduplication, so customer dashboards showed hiring volumes that clients knew were wrong.
Cross-board deduplication to canonical postings with employer resolved to domain, repost linking, and source count retained as a signal.
Reported hiring volumes came into line with what clients recognised from their own recruitment activity.
Posting-count time series were dominated by reposting behaviour and agency duplication, producing spikes that had no relationship to actual headcount growth.
Deduplicated, employer-resolved hiring velocity with agency postings flagged and excludable, and reposts linked rather than counted as new.
The series became stable enough to use as a model input rather than as anecdote.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Deduplication and employer resolution are ongoing modelling work, not a one-time build.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Deduplication gets the posting count right. Employer resolution decides whether you can attribute anything to a company at all — and it is harder than it sounds, because job postings are frequently written to obscure the employer.
We resolve employer to a domain wherever possible, since a domain is the most joinable identifier — it links cleanly to firmographics and to your CRM. Resolution uses posting metadata, ATS board ownership, description signals and name normalisation against registry data.
Agency postings are flagged rather than dropped, because they matter for staffing clients and distort counts for investment clients — opposite requirements from the same field. Where the employer genuinely cannot be resolved, the record says so instead of guessing. A wrongly attributed hiring spike is a far worse outcome than an acknowledged unattributed one.
For teams that also need company-level context, this joins directly to our lead and contact data on the resolved domain.
Job data sits closer to personal data than most scraping categories, and the boundaries deserve stating plainly rather than being buried in a compliance footnote.
Publicly posted vacancies, which are published specifically to be seen and are business information about an employer's activity, not personal data about a person. Employer, title, location, salary, skills and dates all sit comfortably in that category.
You receive a written methodology document identifying every source and what is collected from it, plus a DPA before signature. Job data is one of the areas where legal review is most likely to be thorough, and being able to hand over that document is usually what unblocks the purchase.
Company watchlist or role taxonomy is scoped first, since that determines source selection and deduplication tuning.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect publicly posted job listings from public boards, aggregators and career pages. We do not access professional networks behind a login, candidate databases, applicant records or recruiter portals. Named individuals' contact details are not part of the deliverable. Every engagement includes a written methodology document and a DPA available before signature.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What HR tech, investment and workforce teams ask during evaluation.
We collect job postings that appear on public job boards and company career pages. We do not scrape professional networks behind a login, and we do not collect member profile data from them. That is a firm boundary rather than a capability question, and it is the request we decline most often in this category.
In practice this matters less than clients expect: most roles posted to a professional network are also posted to the employer's ATS-hosted career page, which we treat as the canonical source anyway — and that version is more reliable because it is the employer's own posting rather than a syndicated copy.
Clustering on employer identity, normalised title, location and description similarity, then collapsing to one canonical posting. Our observed ratio is around 3.9 raw postings per real vacancy, and it runs much higher in logistics and retail.
We keep sources_seen as a field rather than discarding it, because posting breadth is itself a signal — a role pushed to eight boards is usually harder to fill. Reposts are linked via repost_of rather than merged, so genuine time-to-fill is measurable instead of reset each time a recruiter refreshes the listing.
About 61% in our overall coverage, and it varies enormously — markets with pay transparency legislation run far higher, and senior roles disclose far less than hourly ones.
We never impute a salary. Ranges are parsed with currency and period where published and left null where not, and we report the publication rate alongside any aggregate. An imputed range inside a compensation benchmark is worse than a visible gap, because it looks like evidence.
Sometimes, and we are explicit about when. Description signals, ATS ownership and posting patterns often identify the employer, and we resolve to a domain where we can. Where the posting genuinely discloses nothing identifiable, the record says unresolved rather than carrying a guess.
Agency postings are flagged either way, because the requirement is opposite depending on who you are: staffing clients want them, investment clients need them excluded so hiring counts are not inflated by agency duplication.
Daily as standard, and several times daily on a defined company watchlist where speed matters — typically for sales triggers or competitive hiring intelligence.
ATS-hosted career pages are the fastest reliable source because postings appear there first, before aggregator syndication. Aggregators can lag the original posting by a day or more, so a watchlist built on career pages materially outperforms one built on boards.
A proxy for it, and the distinction matters. We measure days from first observation to the posting disappearing or being marked closed, with reposts linked so the clock is not reset by a recruiter refresh.
What that does not tell you is whether the role was filled or cancelled — a posting disappearing means it stopped being advertised, which is not the same as a hire. We report it as days_open rather than time-to-fill for exactly that reason, and clients modelling recruitment efficiency should treat it as an upper-bound proxy.
Yes, into structured lists with required and preferred flagged separately, covering technical skills, tools, platforms, certifications, languages and education requirements.
This is one of the more valuable outputs because it makes technology adoption observable: when a platform starts appearing in requirements across an industry, that usually precedes vendor-reported adoption by months. Extraction is imperfect on unusual phrasing, and we report accuracy per skill category rather than a single headline figure.
The United States and United Kingdom are deepest, with dense board ecosystems and widespread ATS-hosted career pages. Germany, France, Netherlands and Nordics are strong. India has enormous volume with high duplication, so deduplication matters more there than anywhere. GCC markets are growing quickly with a concentrated board landscape.
Coverage is thinner where hiring happens through informal channels rather than public postings, and no scraping service can fix that — the postings do not exist. We assess this per market during scoping.
We quote individually. The drivers are source count, whether you need company career pages (more work than boards, and more valuable), geographic scope, deduplication depth and refresh frequency.
A defined company watchlist with career page monitoring at daily refresh sits at the lighter end. Full multi-market board and career page coverage with deep deduplication and skills extraction sits considerably higher. One scoping call, a free pilot on your own watchlist within 48 hours, then a fixed monthly quote. Request a quote.
Send us a company list or a role taxonomy. We return deduplicated, normalised postings with salary and skills extracted within 48 hours.
Free pilot, no card, no obligation. We'll show you the raw-to-canonical ratio on your own sector.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Unlock retail insights with Noon Saudi Arabia Product Data Extraction to track prices, inventory, discounts, and product trends in real time.
Unlock Hertz & Avis Rental Car Data for Dynamic Pricing Intelligence to track rental rates, availability, and market trends in real time.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.