
Data Privacy for AI Deployments in India
Neo Hives IT Solutions· 10 September 2026·12 min read
Most AI projects in India do not die in the build. They die in a legal review, four weeks after everyone agreed it was a good idea, when somebody finally asks what happens to customer data when the model is called. And the reason that question stalls the project is rarely that the answer is bad. It is that nobody wrote it down.
This is a practical guide to the data-protection side of shipping AI in India: what the Digital Personal Data Protection Act framework asks of you in business terms, the five questions any AI vendor should answer in writing, the deletion problem almost nobody designs for, and the architecture choices that make compliance a property of the system rather than a paragraph in a policy.
One caveat, stated plainly: this is engineering guidance, not legal advice. The rules under the Act have been arriving in stages and the detail moves — check the current position with counsel and against the Ministry of Electronics and Information Technology before you rely on any date or figure below.
What the framework asks of you, in business terms
Strip away the drafting and the DPDP Act asks a handful of questions that every AI deployment must be able to answer. Each has a direct engineering consequence, which is the useful way to read it:
| The requirement | What it means for an AI build |
|---|---|
| Consent for a specified purpose | You cannot quietly reuse data collected for one purpose to power something else. Support transcripts gathered to resolve tickets are not automatically available to train or evaluate a model. |
| Notice, in plain language | A person must be told what is collected and why, in clear terms — and the Act requires the notice to be available in English or one of the languages in the Eighth Schedule. If you serve customers in Hindi or Tamil, the notice needs to exist there too. |
| Purpose and storage limitation | Data is kept for as long as the purpose requires, not indefinitely. Call recordings, chat transcripts and model logs all need a retention period set deliberately rather than by default. |
| Correction and erasure rights | A person can ask you to correct or delete their data — and you must be able to do it everywhere it landed, which is the hard part. |
| Reasonable security safeguards | The obligation with the largest penalty attached. Applies to your logs and vector indexes exactly as much as to your primary database. |
| Breach notification | You must be able to detect and report a breach, which presumes logging and monitoring you may not currently have on your AI stack. |
| Children's data | Processing a child's personal data requires verifiable parental consent, and tracking and targeted advertising to children is off the table. A product requirement, not a checkbox. |
| Accountability as data fiduciary | You remain answerable for what your processors do — including the AI provider you call over an API. Their compliance is your problem. |
The dates, and the penalties
The Act commences in phases. The provisions establishing the Data Protection Board came into force on 13 November 2025; a further set follows on 13 November 2026; and the remaining provisions — which include the substantive obligations most businesses are planning for — on 13 May 2027. Two things follow from that timetable. There is time to design properly rather than retrofit. And there is no case for building something now that you already know will be non-compliant then, because the expensive version of this work is doing it twice.
The penalties in the Schedule are assessed by the Board after an inquiry, and the ceilings are not symbolic: up to ₹250 crore for failing to take reasonable security safeguards to prevent a personal data breach, and up to ₹200 crore for failures around children's data. These are maxima rather than tariffs, but the ranking tells you where the regulator's attention sits — security and children, in that order.
On cross-border transfer, the Act takes a permissive shape: transfers outside India are allowed except to countries the central government restricts by notification. That is considerably easier to work with than a residency mandate — but note that sector regulators impose their own rules, and financial and health data often carry stricter localisation requirements than the Act itself. Check your sector before concluding you are free to process anywhere.
The five questions to put to any AI vendor, in writing
This is the list that unblocks legal reviews, because it converts "is this safe?" into answerable specifics. Ask for the answers in the contract, not in an email from a salesperson.
- What data leaves our network on each call, exactly? Not "it's secure" — the actual payload. Usually the prompt, the retrieved passages and the user's question, which together may contain far more personal data than anyone assumed.
- Is anything retained, and for how long? Many providers hold inputs and outputs for a period for abuse monitoring. That is often defensible; it is not something to discover later. Get the number.
- Is our data used to train or improve models? The answer should be a contractual no for business deployments. Verify it in the terms rather than in the marketing.
- Where is it processed, geographically? Region matters for your own commitments to customers and for any sector rules that apply to you.
- How do we delete one person's data from every layer? The question most vendors answer worst, and the one covered next.
The deletion problem nobody designs for
A deletion request sounds like a database operation. In an AI system, personal data has usually spread into at least six places, and only the first is the one people think of:
- Your primary database. Straightforward.
- The vector index. Embeddings of documents containing the person's data are derived data and must be removed too. This means keeping a mapping from source document to every chunk and vector, from the beginning. Retrofitting it means re-indexing everything.
- Prompt and response logs. The debugging logs that made the system supportable are full of personal data. They need retention windows and a way to purge by subject — which in practice means logging an identifier you can search on, rather than a blob of text.
- Caches. Cached answers and cached context can hold personal data for as long as the cache lives. Short, deliberate expiry times are the cheap fix.
- Call recordings and transcripts. The largest volume and the most sensitive, particularly for voice deployments. Recordings accumulate silently and are almost never given a retention policy at launch.
- Evaluation sets. The unglamorous trap. The fixed set of real cases you built to test the system properly is a permanent copy of real customer data sitting outside your main systems. Redact it at the point of creation, or accept that you now have a compliance obligation attached to your test suite.
And one genuinely one-way door: fine-tuning a model on personal data. Once a person's data is in the weights, there is no honest deletion story — only retraining. That single fact is a good reason to prefer retrieval over fine-tuning for anything touching customer data, quite apart from the cost. Retrieval keeps facts in a database you can delete from; fine-tuning bakes them into an artefact you cannot.
Architecture choices that do the compliance work for you
Policy documents do not protect data; design does. Six patterns that reliably reduce the surface area:
- Minimise at the boundary. Before anything leaves your network, ask what the model actually needs. It rarely needs the customer's full name, phone number and address to classify a complaint. Redact or tokenise identifiers on the way out and restore them on the way back — the model works on the shape of the problem, not the identity.
- Separate identity from content. Keep the personal identifiers in your own systems and pass a reference. This is old-fashioned data hygiene and it is unusually effective here.
- Enforce permissions at retrieval time, per user. The single most important control in a knowledge assistant: search returns only what that person may already open. Never rely on instructing the model to keep something secret — if a document reaches the model, treat it as disclosed.
- Set retention windows on everything, at launch. Logs, transcripts, recordings, caches, exports. "We will decide later" becomes years of accumulated liability, and it always accumulates in the least-monitored bucket.
- Log access, not just errors. Who asked what, and which documents were returned. You need this to answer a subject request, to investigate a breach, and to prove the controls work.
- Choose the processing region deliberately, and write it down. It is a one-line configuration decision at the start and a migration afterwards.
For the wider governance frame — risk assessment, documentation, ongoing monitoring — the NIST AI Risk Management Framework is a sensible, non-jurisdictional structure to borrow, and it maps reasonably onto what an Indian data fiduciary needs to be able to show.
Two situations that need extra care
Children. If under-18s use your product — education, gaming, coaching, anything with a student audience — verifiable parental consent is a design constraint that reaches into sign-up, session handling and analytics, and it carries one of the two largest penalty ceilings. It is not something to add in a later sprint; it changes the product.
Employee data. HR assistants over personnel files are one of the most popular internal AI projects and one of the riskiest, because the corpus contains salaries, medical notes, disciplinary records and performance reviews all in the same folder tree. The Act's provisions around employment do not suspend the need for access control. If anything, internal deployments deserve stricter permission testing than customer-facing ones, because the audience is larger and more curious.
What good looks like
- A one-page data flow map: what leaves, to whom, where it is processed, how long it is kept. Written before the build, not after.
- Contract terms with the AI provider covering retention, training use and region.
- Retention periods configured on logs, transcripts, recordings and caches from day one.
- A tested deletion path that covers the vector index, logs, caches and evaluation sets — tested, not assumed.
- Permission tests in the release suite, run with accounts at different access levels every time.
- Consent notices in plain language, available in the languages your customers actually use.
- A named owner for the AI system's data, who is not "the vendor".
Where this goes wrong
- Personal data flows into prompts because nobody looked at what the payload contained.
- Debug logging left at full verbosity in production, quietly building the most sensitive database in the company.
- A model fine-tuned on customer records, making erasure impossible.
- An evaluation set of real customer cases sitting in a repository with no retention policy.
- Call recordings kept forever because deleting them was never anybody's job.
- Vector embeddings not linked back to source documents, so deletion means re-indexing from scratch.
- Permissions enforced by prompt instruction rather than at retrieval.
- Consent for one purpose treated as consent for any future AI use of the same data.
- Compliance treated as documentation, when every one of the failures above is an engineering decision.
How we approach this at Neo Hives IT Solutions
We draw the data flow map before writing code, because it takes an hour and it is the artefact that gets a project through legal review. That means naming what leaves the network on each call, what gets minimised or tokenised first, where processing happens, what is retained and for how long, and how a deletion request would be executed across every layer including the test data.
Then the controls go in with the first version rather than after it: retention windows, per-user permissions at retrieval, access logging, and permission tests in the release suite. It is the same instinct as our testing and quality work — a property you can demonstrate with a test beats a property you assert in a policy. Where the harder question is whether a project should touch personal data at all, that is IT consultation before it is engineering, and often the answer is that a redacted or aggregated version of the data does the job perfectly well. We are an early-stage company, we are not your lawyers, and we will tell you when a question needs one.
Common questions
Does calling a foreign AI provider break Indian law? Not by itself — the Act permits transfers except to countries the government restricts. What matters is that you know where processing happens, that your contract covers retention and training use, and that your own sector rules and customer commitments allow it.
Do we need fresh consent to run AI over data we already hold? It depends on the purpose you collected it for. Using existing support tickets to answer the same customer's support question is usually a continuation of that purpose; using them to build a product feature or a training set is usually not. This is exactly the question to put to counsel, and the reason to keep purpose recorded per dataset.
Does everything need to run on-premises? No, and defaulting to it is expensive. On-premises or private-cloud deployment is the right answer for genuinely sensitive corpora and a costly habit otherwise. Decide it per dataset on sensitivity, not once, on general unease.
Is anonymised data outside all of this? Genuinely anonymised data, yes — but the bar is higher than removing names. Free-text notes, rare combinations of attributes and call recordings are all re-identifiable in practice. Treat "anonymised" as a claim to test rather than a label to apply.
Who is accountable if the AI vendor leaks the data? You are, to your customers and to the regulator, as the fiduciary who chose the processor. That is why the contract terms and the data flow map matter more than the vendor's brand.
Do we need to appoint a data protection officer? The heavier obligations — including a designated officer in India and periodic audits — attach to entities notified as Significant Data Fiduciaries, based on volume and sensitivity of data among other factors. Most small businesses will not be in that category, but if you process large volumes of sensitive personal data it is worth checking rather than assuming.
Where to start
Draw the data flow map for whatever AI feature you are planning, on one page: what leaves your network, to whom, where it is processed, how long anything is kept, and how you would delete one person from all of it. If any box is empty, that is your next task — and it is far cheaper to fill in now than after the system holds eighteen months of transcripts.
Our free AI readiness audit covers this alongside the technical side: what your planned deployment would send where, and what would need to change before it could face a legal review. Or get in touch and tell us what data the idea depends on.