Training data is the most common source of hidden DPDP risk in machine-learning systems. If your training set contains personal data of individuals in India, you need a lawful basis — usually consent under S.6 — for using it to train a model, not just for your original collection purpose. Scraped, purchased, or repurposed datasets are especially high-risk, because the individuals rarely consented to model training and you often cannot prove any lawful basis at all. This checker assesses your training data across source, consent, sensitivity and children's data to flag where your ML pipeline is exposed.
Training sets are where DPDP risk hides. Check whether the personal data used to train your model has a lawful basis — and where it does not.
The core issue is purpose. Under S.6 of the DPDP Act, consent must be specific to the purpose for which personal data is processed. Data a user handed over to receive your service was consented to for that service — not, by default, for training a machine-learning model that may be commercialised, shared, or embedded into products those users never saw. When a team repurposes service data as training data, it is introducing a new processing purpose that usually needs its own lawful basis. Bundling model training into a long, generic privacy policy does not create valid specific consent.
The problem compounds with external data. Scraped web data and purchased datasets are attractive because they are large and cheap, but they carry the weakest consent chain. Individuals whose data was crawled from the public web almost never agreed to model training, and the fact that data is publicly accessible does not, on its own, create a lawful basis under DPDP. Purchased datasets simply move the risk onto you as the data fiduciary using the data. Niti Bharat works with AI teams to trace this consent chain before models are trained, so the training set itself is defensible rather than a latent liability.
The most effective single step is data minimisation. Before training, ask whether the model genuinely needs each field of personal data to learn its task, and strip out what it does not. A model that predicts churn rarely needs raw contact details in its features; a model that classifies documents rarely needs the identities of the people named in them. Reducing the personal data in the training set reduces both your obligations and the harm if the set is ever breached — and security-safeguard failures leading to a breach carry the highest penalties under the Act, up to ₹250 crore.
Beyond minimisation, keep provenance records for every dataset, verify the age boundary to avoid inadvertently training on children's data under S.9, and maintain a link between training data and model versions so you can respond to data principal erasure and correction requests. Niti Bharat's fixed-price DPDP engagements build these controls into the ML lifecycle, giving AI product teams a training-data governance approach that holds up ahead of the May 2027 enforcement date rather than one assembled reactively.
A practical checklist for auditing a machine-learning training set under DPDP — provenance, consent basis, minimisation, sensitive-data handling and children's-data filtering.
One real DPDP development explained in plain English, one practical how-to, one number from our own assessment data. Nothing else — no daily noise, no sales pitch.
No spam. Unsubscribe with one click, anytime.