Programme catalogues
What is actually taught, at what level and in what format.
- Programme title, level and field
- Delivery mode: campus, online, hybrid
- Duration and credit structure
- Intake months and application deadlines
For teams sizing a market nobody has mapped properly.
Half the fee data in this sector lives in PDF prospectuses that nobody has ever structured. That is the part of the service that takes real work.
Free pilot on your own sources, returned in 48 hours. No card, no trial clock — and you keep the sample data either way.
Last verified 5 August 2026 by the Actowiz Solutions Data Engineering team.
Education and EdTech data spans two related universes. The institutional side covers universities, colleges and schools: their programme catalogues, tuition and fee schedules, entry requirements, accreditation, intake cycles and published outcome statistics. The EdTech side covers digital learning platforms: course listings, pricing and subscription tiers, enrolment counts, ratings and instructor data.
Both are nominally public. Neither is remotely comparable without substantial normalisation work.
Actowiz handles this with a document-processing layer — including OCR for scanned schedules — feeding a normalisation model that resolves pricing unit, student category and academic year, then retains a reference to the exact source page for every fee. When a figure looks wrong, your analyst can open the source PDF and check it rather than filing a support ticket.
Institutional and EdTech data share one schema where fields overlap, so a single query can compare a campus master's against an online equivalent.
What is actually taught, at what level and in what format.
The full cost picture, normalised to a comparable unit.
The organisational context that makes programmes comparable.
The eligibility layer that drives student-recruitment matching.
The online learning market, refreshed far more frequently.
The signals prospective students and investors weigh most heavily.
Institution mapping, programme normalisation and fee extraction including documents.
Every engagement delivers a documented schema. These are the core fields; the full dictionary is agreed during scoping.
| Field | Type | What it captures | Refresh |
|---|---|---|---|
programme_id / institution |
string | Stable programme identifier and normalised institution name | Termly |
level / mode |
enum | Certificate, diploma, bachelors, masters, doctoral; campus, online or hybrid | Termly |
duration_months / credits |
int | Programme length normalised to months, plus credit total where stated | Termly |
tuition_domestic / international |
decimal | Tuition by student category in local currency | Termly |
tuition_unit |
enum | per_year, per_term, per_credit, per_module or full_programme | Termly |
additional_fees |
array | Bench, lab, registration, deposit and insurance charges itemised | Termly |
fee_source_ref |
string | Reference to the exact page or PDF location the fee was extracted from | Termly |
entry_requirements |
object | Grade thresholds, English test scores, prerequisites, experience requirements | Termly |
accredited / accreditor |
boolean / string | Accreditation status and the accrediting body where published | Annual |
platform_price / tier |
decimal / string | EdTech course price and subscription tier where applicable | Weekly |
enrolment_count / rating |
int / float | Publicly displayed enrolment volume and average rating | Weekly |
Where an institution publishes fees only as a scanned PDF, we OCR it and mark the field with an extraction-method flag so you know which values came from document processing rather than structured markup.
Coverage is strongest where institutions publish in English and in structured formats, but PDF and OCR handling extends it considerably beyond that.
Non-English institutional sites are supported with in-language extraction and optional translation of programme titles and descriptions. Request a source we don't list →
We deliver into 40+ countries. These are the markets where this particular service is requested most, and the reason demand concentrates there.
| Market | Why demand concentrates here |
|---|---|
| United States | Tuition and programme data is fragmented across thousands of institutions. |
| United Kingdom & Australia | Major international student destinations; fee transparency drives recruitment. |
| India | 90,000+ institutions and the fastest-growing EdTech competitive landscape. |
| United Arab Emirates & Singapore | Dense international and transnational education markets. |
We run production collection across 40+ countries. Coverage depth varies by market and by source, so we confirm what is actually available for your specific markets during scoping rather than claiming uniform global coverage. Ask about a market we don't list →
Four distinct buyer groups with quite different priorities — pricing, investment, recruitment and policy.
You are pricing courses against competitors whose prices change weekly, and manual checks cover a fraction of the catalogue.
Weekly competitor course pricing across major platforms, with tier structure, enrolment volume and rating, so pricing decisions reflect current market position.
Price competitiveness
Education market sizing relies on vendor reports with opaque methodology, which makes diligence on a specific segment or geography unreliable.
Bottom-up programme and institution universes with fee data, so market size is built from observable prices and provision rather than a top-down estimate.
Sizing confidence
Matching students to programmes requires current fees, entry requirements and deadlines across hundreds of institutions, maintained manually in spreadsheets.
Normalised programme data with fees, requirements and intake dates refreshed termly, delivered into your matching platform or CRM.
Placement conversion
You need to know how your fees and programme portfolio compare to your competitor set, but building that view takes an analyst weeks each cycle.
Fee and portfolio benchmarking against a defined peer group, refreshed each cycle, with pricing unit normalised so comparisons are valid.
Portfolio competitiveness
Cross-country education cost and provision comparison requires normalising data published in dozens of formats and currencies.
Harmonised cross-border datasets with pricing unit, currency and student category normalised, plus source references for every figure.
Research reproducibility
Selecting training providers means comparing hundreds of courses on price, format and rating with no consolidated view.
Consolidated course catalogues across professional learning platforms with pricing, duration, rating and enrolment volume in one schema.
Cost per learner
Four patterns, with measured outcomes.
Weekly extraction across competing platforms captures course-level pricing, subscription tier structure, discounting patterns, enrolment volume and ratings. Because pricing unit and tier are normalised, a subscription-bundled course can be compared to a standalone purchase on a consistent basis.
Outcome: Pricing decisions informed by current competitor positioning rather than a quarterly manual review.
For a defined segment — say postgraduate business programmes in Western Europe — we build the complete programme universe with fees, modes and institution profiles. Market size is then computed from observable provision and pricing rather than extrapolated from a top-down report.
Outcome: Diligence and market entry decisions grounded in an auditable universe with source-linked figures.
Recruitment agencies receive normalised programme data with fees by student category, entry requirements, English test thresholds, intake months and deadlines, refreshed termly and delivered into their matching platform.
Outcome: Advisors working from current requirements and fees rather than spreadsheets that go stale within a term.
An institution's fee schedule is compared against a defined peer group, programme by programme, with pricing unit normalised and additional fees included so the comparison reflects total cost rather than headline tuition.
Outcome: Fee-setting decisions supported by valid peer comparison rather than headline figures that aren't like-for-like.
Clients rarely permit naming. These are real engagement shapes with identifying detail removed, so you can judge whether the work resembles your situation.
The team needed programme and fee benchmarks across thousands of institutions, but most fee schedules existed only as PDFs with inconsistent structures.
PDF and OCR fee extraction with tuition units normalised across per-year, per-term and per-credit, retaining a source reference for every figure.
A structured fee benchmark replaced a spreadsheet built from manual sampling.
New programmes, fee changes and entry requirement updates at competitor institutions were tracked manually and inconsistently.
Termly institution and programme monitoring with change detection on fees, entry requirements and programme launches.
Programme and fee changes surfaced within one collection cycle instead of by chance.
Examples are anonymised at client request. Named references are available on request under NDA. See published case studies →
Before you commit to anything, we run this service against your own sources and send you the output. If the coverage isn't there, the sample will show you that too — which is the point. We would rather lose the deal at the pilot than at month three.
Same collection pipeline and same QA underneath. The difference is who holds the schedule and how the data reaches you.
We own the collection, the QA and the delivery. You receive clean data on a schedule and never touch a scraper.
Best fit: Teams who need the data, not the infrastructure.
The same collection pipeline exposed as an authenticated REST endpoint your systems query directly.
Best fit: Product and engineering teams building on live data.
A defined pull for a specific question — market sizing, diligence, a pitch, a one-off audit.
Best fit: Research, strategy and diligence work with a deadline.
Every engagement is quoted individually, because the honest answer depends on your scope: how many sources, how many records, how often, and how the data reaches you. We scope it with you, run a free pilot on your own sources, and then quote a fixed monthly figure — no per-request metering and no overage billing when volumes move. Request a quote and you will have a number after one call.
Document-based fee extraction is where in-house attempts in this sector usually stall.
| Consideration | In-house scraping team | Generic proxy / DIY tool | Actowiz managed feed |
|---|---|---|---|
| Time to first usable data | 6–12 weeks of engineering before anything is trustworthy | Days, but output needs manual cleanup before use | Free pilot in 48 hours, production in 5–10 business days |
| Who fixes it when a source changes | Your engineers, at the cost of their roadmap | You do — tools report failures, they don't resolve them | We do, same business day, inside the retainer |
| Data quality assurance | Whatever your team has time to build | None beyond HTTP success | Schema validation plus sampled human QA on every run |
| Compliance documentation | Rarely produced, then requested urgently by legal | Not provided; terms risk sits with you | Sources, method and lawful basis documented for review |
| Accountability | Distributed across a team with other priorities | A support ticket queue | A named engineer and an account owner |
| True annual cost | Engineer salaries, proxies, hosting, ongoing maintenance | Low licence fee plus significant hidden analyst time | One fixed monthly retainer, quoted after scoping |
Retail pricing has one enormous advantage: a product has a price, and that price is a single number attached to a purchasable unit. Education pricing has none of that structure.
A master's degree might be advertised at £14,500 — but that figure could be per year on a two-year programme, per full programme, or the domestic rate where you needed the international one. Add a £1,200 bench fee published on a separate page, a mandatory health insurance charge for international students, and a deposit deducted from year one, and the advertised number bears limited relationship to the amount a student actually pays.
We keep the original extracted string alongside every normalised value. When your analyst disagrees with our interpretation, they can see exactly what the source said and apply their own judgement, rather than trusting a black box.
Education data has become a significant input for AI teams — both for building learner-facing products and for market intelligence models. The requirements differ meaningfully from a BI feed.
If you are building a learner-facing product, be aware that fee accuracy carries real reputational stakes — a quoted figure that turns out wrong is worse than no figure. Our source-reference field exists partly so products can link users to the authoritative page rather than asking them to trust an extracted number.
Institution and platform lists are scoped first, with an honest coverage assessment before contracting.
You send us target sites, regions, SKUs or keywords. We return a field-level schema proposal, coverage estimate and refresh recommendation — usually within two working days.
We extract a real sample from your actual targets so you can inspect field fill rates, edge cases and match quality before any commitment.
Our engineers build extractors, then wire validation rules: type checks, range checks, duplicate detection and golden-record comparison against a manually verified subset.
Feeds run at your chosen cadence and land in the warehouse or bucket you already use. Schema changes are versioned and announced before they ship.
We watch coverage drift, fill rates and source changes daily. A named engineer owns your account, and layout breaks are fixed by us — not queued for you.
JSON, JSONL, CSV, Parquet or XLSX, delivered to Amazon S3, Google Cloud Storage, Azure Blob, SFTP, Snowflake, BigQuery, Databricks or a REST/GraphQL endpoint. Webhooks fire on completion, and every batch ships with a manifest containing row counts, schema version and QA results so your pipeline can fail loudly instead of silently ingesting a bad file.
We collect only publicly accessible information, respect robots directives and rate limits, never bypass authentication or paywalls, and never scrape personal data outside a documented lawful basis. Each engagement includes a written collection methodology, source list and retention policy your legal and procurement teams can review before signature.
These are contractual, not marketing copy. They appear in the engagement document.
| Commitment | What we hold ourselves to |
|---|---|
| Pilot turnaround | A real sample from your own sources within 48 hours of scoping, at no cost. |
| Go-live | Production collection running within 5–10 business days of sign-off. |
| Delivery punctuality | 99.5% on-schedule delivery, measured monthly and reported to you. |
| Breakage response | Source layout changes triaged same business day; critical sources inside 4 hours. |
| Data quality | Schema validation on every run plus sampled human QA before any delivery leaves us. |
| Escalation | A named engineer and an account owner, not a shared ticket queue. |
| Change requests | Field additions and source changes handled inside the retainer, not re-quoted. |
| Exit | Your historical data exported in full on request. No lock-in, no export fee. |
Plain definitions of the terms used on this page, so procurement and legal reviewers are working from the same vocabulary as your data team.
What buyers ask during evaluation.
Yes, and this is essential rather than optional in this category. A large share of institutions publish fee schedules only as PDFs, frequently as scanned tables. We run a document-processing pipeline with OCR for scanned material, then normalise the extracted values into our schema.
Every fee carries an extraction-method flag, so you know which values came from structured markup and which came from OCR. OCR-derived values are held to a higher review threshold, and low-confidence extractions go to human verification before delivery.
We resolve the pricing unit explicitly and store it in a dedicated tuition_unit field with values like per_year, per_term, per_credit or full_programme. Where the institution states it, we use their statement; where it must be inferred from duration and credit structure, we infer it and flag the inference.
Where the unit genuinely cannot be determined with confidence, we mark it unresolved rather than guessing. A wrong unit is worse than a missing one, because it produces comparisons that look valid and aren't — a per-year fee compared against a full-programme fee understates cost by a factor of two or three.
Yes. We extract in-language and optionally supply translated programme titles and descriptions alongside the originals. Coverage is strongest in Western European languages, and we handle Hindi, Arabic and East Asian scripts with reasonable reliability.
Extraction quality is generally lower on non-English sites, mainly because page structures vary more and fee presentation conventions differ. We give you honest per-region coverage expectations during scoping rather than a uniform promise followed by a file full of nulls.
On different cycles depending on the field. Tuition and fees change annually, usually announced several months before the academic year. Programme catalogues change termly as courses are added and retired. Entry requirements change annually. EdTech platform pricing changes weekly or faster, with frequent promotional discounting.
We refresh institutional data termly and EdTech pricing weekly, which matches how these actually move. Refreshing university fees weekly would generate identical records at real cost; refreshing EdTech pricing termly would miss most of the pricing behaviour.
Only what institutions and platforms publish publicly, which varies enormously. Some jurisdictions mandate publication of completion rates, employment outcomes and salary data; others publish nothing. EdTech platforms often display enrolment counts and ratings openly.
We extract what is published and clearly mark what is absent. We do not model or estimate enrolment figures — where a number does not exist publicly, we supply no number. Vendors who fill those gaps with estimates rarely disclose the method, and the resulting figures find their way into market sizing where they cause real damage.
Yes, and clients do exactly that — but with a caveat worth stating. Fee accuracy carries reputational stakes in student-facing products: a quoted figure that turns out to be wrong, or the wrong student category, damages trust immediately and is hard to recover.
For this reason we recommend surfacing our fee_source_ref field in your product so users can click through to the authoritative institutional page. That way your product provides the comparison and the institution provides the confirmation, which is the right division of responsibility.
We extract scholarship and bursary listings where institutions publish them, including eligibility criteria, award value and application deadlines where stated. Coverage is inconsistent because publication practice is inconsistent — some institutions maintain structured scholarship databases and others mention funding in prose on a single page.
Scholarship data decays quickly since deadlines pass and awards are withdrawn. We capture the stated deadline and our capture date so your product can suppress expired listings rather than showing a student an opportunity that closed months ago.
We quote every education data engagement individually, because a real number depends on scope: source count, record volume, refresh frequency and delivery method. Anyone quoting you a price before understanding those four things is guessing.
Institution count and how much of the fee data sits in PDFs are the main drivers, since document extraction is more labour-intensive than page collection.
The process is short: one scoping call, a free pilot on your own sources within 48 hours, then a fixed monthly quote. No per-request metering, no overage billing, and field or source additions are handled inside the retainer rather than re-quoted. Request a quote.
Yes, and it is one of the strongest use cases here because no comparable public dataset exists. For a defined segment — level, field, geography, delivery mode — we assemble the complete provision universe with fees, institution profiles and mode, then deduplicate across institutional sites, aggregators and recruitment portals.
We will also tell you honestly where the universe is incomplete. In some markets institutional publication is patchy enough that a genuinely complete list is not achievable, and knowing that before you build a model on it is more valuable than a confident number that isn't.
Send us institutions or a target region. We return normalised programme and fee records with source references within 48 hours, at no cost.
Free pilot, no obligation, no card. You'll have a fixed monthly quote after one scoping call.Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
The EU AI Act impact on web scraping & AI training data GPAI transparency, copyright reservations, prohibited practices & a compliance checklist from Actowiz.
How a B2B supplier replaced manual tender-portal checking with an automated, filtered feed of relevant government tenders from GeM and CPP/eProcure never missing a bid deadline again.
Actowiz Solutions tracks post–World Cup 2026 travel pricing — hotel ADR & airfare normalization across host cities, event-premium decay data & lessons for travel teams.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.