AI and ML companies face acute DPDP Act 2023 exposure because models are trained on large datasets that often contain personal data. Under the Act, using personal data to train or fine-tune models needs a lawful basis (usually consent or a recognised legitimate use), purpose limitation applies, and scraping personal data from the web is high-risk. This guide and checker help AI/ML teams in India size and reduce their data-protection risk.
Training data, web scraping, consent and model risk — how India's DPDP Act 2023 applies to AI and ML teams.
Generative and predictive models are only as good as their training data, and that data frequently includes personal information — user content, support transcripts, scraped text, purchased datasets. The DPDP Act 2023 does not exempt AI: if personal data of individuals in India is processed, the Act's consent, purpose-limitation, security and rights obligations all apply.
The hardest issues are lawful basis for training, provenance of scraped or purchased data, and honouring erasure once data is baked into a model. Teams that design for these from the start avoid expensive retraining and regulatory risk later.
Lawful-basis guidance for training data, a dataset provenance template, and a model DPIA starter.
One real DPDP development explained in plain English, one practical how-to, one number from our own assessment data. Nothing else — no daily noise, no sales pitch.
No spam. Unsubscribe with one click, anytime.