Every legally marketable housing project in India must be registered with a state RERA authority, and those registrations are public. That makes RERA the closest thing to a census of Indian residential supply. The obstacle is not access — it is that thirty-plus authorities run thirty-plus independently built portals with inconsistent field naming and genuinely different disclosure depth. This guide covers what is typically published, what varies by state, and how to build a comparable multi-state dataset without drawing false conclusions from missing fields.
The Real Estate (Regulation and Development) Act, 2016 requires promoters to register a project with their state's regulatory authority before advertising, marketing or selling it. Registration carries mandatory disclosure — project details, promoter identity, approved plans, unit inventory, declared timelines — and the authority publishes it.
For anyone analysing Indian housing, this is unusual and valuable. Most real-estate market data is derived from listings, which show what is being marketed. RERA shows what has been approved. The gap between those two is where a large share of interesting analysis lives.
Who uses it, and for what:
| User | Use case |
|---|---|
| Developers | Competitive supply mapping, launch timing, locality saturation |
| Investors and lenders | Underwriting residential exposure, sponsor track-record checks |
| Consultancies and research | Market sizing, absorption analysis, affordable-housing studies |
| Proptech platforms | Verified project inventory, listing validation |
| Policy researchers | Delivery against affordable-housing targets, completion-delay analysis |
Across authorities, the recurring field groups are:
This is the section that determines whether a multi-state dataset is analytically sound.
The single most valuable field for absorption analysis — booked versus unbooked units — is published by some authorities and not others. Where it exists, it typically comes from quarterly progress filings rather than the registration record, which means it updates on a quarterly rhythm and lags reality.
Consequence: you cannot compute a comparable absorption rate across all states. You can compute it for the states that disclose, and you must be explicit that the rest are not in the denominator.
What one authority calls "Project Status" another splits between a status field and an extended-completion-date field, from which status must be inferred. Unit configuration may be a structured table in one state and free text in another. Area figures may be filed in square metres or square feet, inconsistently within the same portal.
Some authorities retain deregistered, lapsed and completed projects in the public registry. Others surface only currently valid registrations. If a state prunes its registry, you cannot reconstruct historical supply from it — you can only build a forward time series from the date you started collecting.
This has a hard consequence: start collecting before you need the history, because for pruning states it is unrecoverable.
Large projects are often registered as separate phases with separate registration numbers. Some authorities link phases explicitly; others do not. Without phase linkage, one 2,000-unit township can appear as eight unrelated projects and inflate your project count while deflating average project size.
Some analyses need registration data from land-records systems — IGR Maharashtra being the common example — alongside RERA. These are separate systems with separate structures, and joining them requires an address or survey-number match that is rarely clean.
Distinguish "not published by this authority" from "published but empty."
This sounds like a minor schema detail. It is the difference between a usable multi-state dataset and one that generates confidently wrong conclusions.
If a state does not publish unbooked-unit counts and your schema leaves that field blank, a downstream analyst sees a blank and reads it as zero — concluding that projects in that state are fully booked. It is a plausible, catastrophic misreading that no data-quality check catches, because the field is legitimately empty.
The fix is a per-field availability marker: this authority does not disclose this field / this authority discloses it and this project filed nothing / this authority discloses it and the value is zero. Three distinct states, and collapsing them loses the distinction permanently.
RERA registration data is published as a statutory public-disclosure requirement. Publication is the purpose of the registry — the Act exists to give buyers visibility into project status and promoter accountability.
The practical boundaries for collection: only publicly accessible logged-out pages, no authenticated or restricted sections, no attempt to access filings the authority has not published. Promoter contact details published on the registry are public business information, but if you intend to use them for outreach, India's data-protection framework and telecom regulations apply to that use independently of how the data was obtained.
Confirm your specific use case with legal counsel. What is publicly published and what you may do with it are two separate questions.
A first phase that produces something useful:
That disclosure map is what lets you scope a wider program honestly. Without it, any multi-state estimate is guesswork about data you have not seen.
Actowiz Solutions has delivered RERA and property-registry extraction across Maharashtra (including RERA Maharashtra district-level project data and IGR Maharashtra land records for Thane, Pune, Mumbai and Navi Mumbai), Telangana project and agent registries, and Gujarat, alongside international property-portal work including Zillow across approximately 3,000 Florida and Texas ZIP codes.
RERA portals publish project registration data as a statutory public-disclosure obligation — public availability is the point of the registry. Collecting publicly accessible published pages is generally accepted practice; authenticated or restricted areas are not in scope. Separately, how you use published promoter contact details is governed by data-protection and telecom rules regardless of collection method. Confirm your use case with legal counsel.
Disclosure varies by authority, and typically comes from quarterly progress filings rather than the registration record. Some states publish it, others do not. Any multi-state dataset should carry a per-field availability marker so absent disclosure is not misread as zero.
Monthly with adhoc refreshes fits most use cases. Authorities update on quarterly filing cycles, so higher frequency largely re-reads unchanged records without adding signal.
Yes, and it is one of the more useful joins available in Indian real estate. RERA shows approved inventory; listing portals show what is actively marketed. The gap between them supports absorption and marketing-intensity analysis.
There is no structural limit, but each authority needs its own extraction profile because no two portals share a structure. Coverage expands state by state rather than all at once.
Cross-state normalization, specifically distinguishing fields an authority does not publish from fields published as empty. Getting that wrong produces confidently incorrect conclusions that no data-quality check will flag.
Only where the authority retains lapsed and completed registrations in its public registry. States that prune to currently-valid registrations only cannot be backfilled — which is a strong argument for starting collection before the history is needed.
You can also reach us for all your mobile app scraping, data collection, web scraping , and instant data scraper service requirements!
Our web scraping expertise is relied on by 4,000+ global enterprises including Zomato, Tata Consumer, Subway, and Expedia — helping them turn web data into growth.
Watch how businesses like yours are using Actowiz data to drive growth.
From Zomato to Expedia — see why global leaders trust us with their data.
Backed by automation, data volume, and enterprise-grade scale — we help businesses from startups to Fortune 500s extract competitive insights across the USA, UK, UAE, and beyond.
We partner with agencies, system integrators, and technology platforms to deliver end-to-end solutions across the retail and digital shelf ecosystem.
Unlock retail insights with Noon Saudi Arabia Product Data Extraction to track prices, inventory, discounts, and product trends in real time.
Unlock Hertz & Avis Rental Car Data for Dynamic Pricing Intelligence to track rental rates, availability, and market trends in real time.
Brazil Car Rental Pricing Intelligence Report 2026 reveals rental price trends, market shifts, competitor rates, and opportunities for smarter pricing.
Whether you're a startup or a Fortune 500 — we have the right plan for your data needs.