Whether the DPDP Act applies to synthetic data in India depends entirely on re-identifiability. The DPDP Act 2023 governs digital personal data of identifiable individuals — so synthetic data that cannot be linked back to any real person generally falls outside its scope. But 'synthetic' is not a magic exemption: if your generator memorised or leaked real records, if the synthetic set can be reverse-engineered to re-identify individuals, or if it was derived from personal data without a lawful basis for that derivation, the Act can still apply. This tool checks whether your synthetic data is genuinely out of scope or still carries DPDP obligations.
Synthetic data is only outside the DPDP Act if no real person can be identified from it. Check whether yours genuinely qualifies — or still carries obligations.
There is no blanket synthetic-data exemption in the DPDP Act 2023. The Act applies to digital personal data — information about an identifiable individual. So the real question is never 'is it labelled synthetic?' but 'can any real person be identified or re-identified from it?'. If the answer is genuinely no, the data is not personal data and falls outside the Act, in the same way as robustly anonymised data. If the answer is yes, or even 'possibly', the synthetic label offers no protection and every DPDP duty applies.
This matters because synthetic data is often generated by training a model on real records, and generative models are known to memorise and reproduce rare or outlier examples. A synthetic dataset that contains a near-copy of a real patient, customer or employee is not anonymous — it is personal data wearing a costume. Niti Bharat helps Indian AI and analytics teams test that boundary properly, so they know which of their synthetic datasets are truly out of scope and which still need to be governed as personal data.
Synthetic data creates DPDP obligations in three main situations. First, when the synthetic set is re-identifiable — outliers, rare combinations or memorised records let someone link a row back to a real person. Second, when the act of generating it counted as processing personal data without a lawful basis — using real customer records to train a generator is itself processing, and needs consent or a documented legitimate use for that purpose. Third, when synthetic and real data are mixed, which quietly pulls the whole dataset back into scope. In each case, penalties for general obligation failures reach up to ₹50 crore, and far more where a security failure leads to a breach.
Because the DPDP Rules 2025 were notified in November 2025 with enforcement expected around May 2027, teams betting their compliance strategy on 'we only use synthetic data' should validate that claim now. Niti Bharat's fixed-price DPDP engagements (₹75K–₹3.2L) include a synthetic-data classification review — a documented re-identification assessment that tells you exactly which datasets are safe to treat as non-personal.
A re-identification test checklist, a synthetic-vs-personal decision tree, and a documentation template to evidence the non-personal status of qualifying datasets.
One real DPDP development explained in plain English, one practical how-to, one number from our own assessment data. Nothing else — no daily noise, no sales pitch.
No spam. Unsubscribe with one click, anytime.