Protecting AI Training Datasets as Trade Secrets: Trade Secret Law, Copyright, and the Unsettled Indian IPR Regime : Author: Shigarfa Shamshad

India has no dedicated trade secrets statute, yet an enormous share of commercial value in artificial intelligence now sits precisely in the kind of asset that regime was never built for: curated, weighted, and painstakingly assembled training data. This piece examines how Indian law currently protects — or fails to protect — AI training datasets, tracing the doctrine from the common law action for breach of confidence through to the Copyright Act 1957 and the Digital Personal Data Protection Act 2023 (DPDPA).

ARTICLE

Shigarfa Shamshad

9/20/2026

Abstract

India has no dedicated trade secrets statute, yet an enormous share of commercial value in artificial intelligence now sits precisely in the kind of asset that regime was never built for: curated, weighted, and painstakingly assembled training data. This piece examines how Indian law currently protects — or fails to protect — AI training datasets, tracing the doctrine from the common law action for breach of confidence through to the Copyright Act 1957 and the Digital Personal Data Protection Act 2023 (DPDPA). It pays particular attention to the complications that arise when training corpora are built from literary and journalistic works, using the Delhi High Court's 2026 ruling in ANI Media Pvt Ltd v OpenAI OpCo LLC as the central case study, and argues that the current patchwork leaves both AI developers and rights-holders in a state of structural uncertainty that only legislative intervention can resolve.

I. Introduction

A trained AI model is, in a meaningful sense, the residue of its training data. The choices behind a dataset — what to include, what to exclude, how to clean and weight it — are frequently what separates a mediocre model from a market-leading one. Firms understandably want to protect that investment the way they would protect a client list or a manufacturing process: as a trade secret. The difficulty is that Indian law has never had a trade secrets statute to begin with, and the doctrine that has grown up in its place — equitable, contract-based, and built for a world of departing employees and leaked client databases — sits awkwardly on top of something as diffuse and legally ambiguous as a billion-document training corpus.[1] Layered onto this is a second, separate problem: much of that training data is not the developer's own information at all, but someone else's copyrighted expression, frequently drawn from news reporting, books, and other literary works. The result is a genuinely fragmented regime, spanning contract law, equity, copyright, and now data protection statute, none of which was drafted with generative AI in mind.[2]

II. Trade Secret Protection in India: An Equitable Patchwork

Unlike the United States (Defend Trade Secrets Act 1996) or the European Union (Trade Secrets Directive 2016/943), India has no standalone legislation defining a trade secret or prescribing a cause of action for misappropriation.[3] Protection instead rests on the equitable doctrine of breach of confidence, supplemented by contract — principally confidentiality clauses under the Indian Contract Act 1872 and, in employment contexts, restrictive covenants read subject to Section 27 of that Act, which renders agreements in restraint of trade void.[4]

The doctrine has been developed almost entirely through a handful of recurring fact patterns: a former employee taking a client database (Burlington Home Shopping Pvt Ltd v Rajnish Chibber[5]), a lawyer departing with case files and precedents (Diljeet Titus v Alfred A Adebare[6]), or an employee accused of carrying confidential business information to a competitor (American Express Bank Ltd v Ms Priya Puri[7]). Courts in these cases have generally required the claimant to show that the information was confidential in nature, imparted in circumstances importing an obligation of confidence, and used or disclosed without authorisation — the classic three-part test derived from English authority in Coco v A N Clark (Engineers) Ltd, which Indian courts have consistently adopted.[8]

The problem for AI training data is that this doctrine was built around discrete, identifiable pieces of information — a formula, a list, a set of drawings — held by a small number of people under a clear duty of confidence. A modern training corpus is neither discrete nor easily bounded: it may run to billions of tokens drawn from thousands of sources, assembled and re-weighted iteratively, and handled by internal teams, external annotation vendors, and cloud infrastructure providers all at once.[9] Recent Indian scholarship has flagged exactly this mismatch, describing the “composited value” of a large aggregated dataset as lacking the identifiability and durability that conventional breach-of-confidence doctrine assumes.[10] In practice, what actually gets protected is not the dataset as such but the access to it — through non-disclosure agreements at every external boundary (annotation vendors, evaluation contractors, red-teaming partners), IP-assignment clauses in employment contracts, and technical access controls and audit logs that stand in as evidence of “reasonable measures” of secrecy, a factor courts and international commentators treat as central to any confidentiality claim.[11]

III. Complications for Literary and Journalistic Works

The more acute legal difficulty, however, is not about protecting a developer's own dataset — it is about the developer's use of someone else's literary or journalistic output within it. Large language models are trained substantially on scraped web content, and a considerable proportion of that content is copyrighted expression: news articles, books, essays, and other literary works within the meaning of Section 2(o) of the Copyright Act 1957.

This tension came to a head in ANI Media Pvt Ltd v OpenAI OpCo LLC, filed before the Delhi High Court in late 2024, in which the news agency ANI alleged that OpenAI had scraped, stored, and used its copyrighted news reports to train ChatGPT without licence, and sought an interim injunction.[12] On 24 July 2026, Justice Amit Bansal declined to grant interim relief, holding — at the prima facie stage — that OpenAI's storage and use of ANI's content for training purposes fell within the fair dealing exception under Section 52(1)(a)(i) of the Copyright Act, which permits fair dealing with a literary work for the purposes of private or personal use, including research.[13] The Court reasoned that the training environment was a closed, non-public one and therefore analogous to private research use; that ChatGPT's outputs were transformative rather than substitutive of ANI's own reporting; and that copyright protects the expression of facts rather than the underlying facts themselves, meaning the threshold for infringement in respect of news content is comparatively high.[14] Notably, the Court also rejected the argument that Section 52(1)(a) is confined to non-commercial use, observing that Parliament had expressly imposed such a limitation elsewhere in the Act but chosen not to do so here.[15]

The ruling is significant but deliberately narrow: it is an interim finding, not a final determination on the merits, and the Court itself left open whether the same reasoning would extend to more expressive literary forms — music, film, or fiction — where the line between fact and protected expression is far less clean than it is in factual news reporting.[16] This is the crux of the “literary complication”: Section 52 of the Copyright Act contains no statutory text-and-data-mining exception of the kind found in the EU's DSM Directive, nor a flexible fair-use standard of the kind applied in the United States. It is a closed, enumerated list of permitted purposes, and whether “training an AI model” can be read into that list — as the Delhi High Court effectively did in ANI — remains a matter of judicial interpretation rather than settled statutory entitlement, leaving authors, publishers, and AI developers alike without clear ex ante guidance.[17] Separately, and largely for want of creative selectivity in assembling data that aims at comprehensiveness rather than originality, the compiled dataset itself is generally not considered an original literary work capable of copyright protection under Indian law, foreclosing that route as a means of protecting the dataset as an asset.[18]

IV. The Data Protection Overlay: DPDPA 2023

A third layer of complication arises where training data includes personal information, which is common wherever web-scraped or user-generated content forms part of a corpus. The Digital Personal Data Protection Act 2023 grants every data principal a right to obtain information about what personal data a fiduciary holds on them and how it is processed, echoing the access-request architecture found in comparable regimes elsewhere.[19] This creates a direct structural tension with the confidentiality interest AI developers assert over their datasets: a request under the DPDPA effectively asks a company to disclose whether, and how, an individual's data sits inside a corpus the company otherwise wishes to keep secret as a trade asset.[20] Indian commentary has observed that this conflict is largely unresolved because the doctrine of breach of confidence presupposes a relationship of confidence between the parties, which simply does not exist between an AI developer and a data subject whose information was scraped from the open web with no direct dealing between them at all.[21] The DPDPA's consent-centric design, built around individual, purpose-specific consent, also sits uneasily with training practices operating at a scale where obtaining or even tracing consent from millions of individual data principals is functionally impracticable.[22]

V. Toward Reform

None of this is an accident of drafting so much as a symptom of timing: the Copyright Act dates to 1957, the confidence doctrine derives from mid-twentieth-century English equity, and the DPDPA was drafted principally with consumer data flows in mind rather than large-scale model training. The Draft National IPR Policy of 2016 gestured toward the need for a standalone trade secrets framework in India but no such legislation has since been enacted, and current DPIIT-level discussion of generative AI and copyright remains at the working-paper stage rather than producing settled statutory change.[23] Until Parliament or the courts on a fuller record address the question squarely, Indian AI developers are left assembling protection out of the tools available — contract, equity, and a favourable but interim reading of fair dealing — while literary rights-holders are left litigating case by case, with the ANI appeal now pending before the Delhi High Court as the next real test of how far that interim reasoning will hold.[24]

VI. Conclusion

The honest position, as matters stand in late 2026, is that AI training datasets in India are protected only imperfectly and indirectly: as confidential information through contract and equity where the data is the developer's own, and as fair-dealt material where it is drawn from someone else's copyrighted literary work, subject to a judicial interpretation that has not yet been tested to finality. For a law student examining this space, the more interesting doctrinal question is not who currently wins under the existing rules, but how ill-suited those rules — built for departing employees and printed newspapers — are to the genuinely novel legal object a training corpus represents.

References

1. Ramendra Mandal, "Privacy, Copyright, and Trade Secret: The Growing Conflict Between Personal Data Rights and AI Training in India" (NASSCOM Community, 2026).

2. "Intellectual Property and Personal Data in AI Datasets Under India's DPDP Act 2023" (2026) IPR Trends.

3. Intepat, "Intellectual Property Law for AI in India 2026" (FAQ: "Does India have a trade-secret statute that protects AI training data and model weights?").

4. Indian Contract Act 1872, s 27.

5. Burlington Home Shopping Pvt Ltd v Rajnish Chibber 1995 (61) DLT 6 (Delhi HC).

6. Diljeet Titus v Alfred A Adebare 2006 (32) PTC 609 (Delhi HC).

7. American Express Bank Ltd v Ms Priya Puri 2006 (110) FLR 1061 (Delhi HC).

8. Coco v A N Clark (Engineers) Ltd [1968] FSR 415, as applied in Indian breach-of-confidence jurisprudence.

9. Mandal (n 1).

10. "Intellectual Property and Personal Data in AI Datasets" (n 2), Abstract.

11. Intepat (n 3).

12. "ANI Media v OpenAI: The Delhi High Court Weighs In" (Mondaq India, 2026).

13. ANI Media Pvt Ltd v OpenAI OpCo LLC, Delhi High Court, order dated 24 July 2026 (Bansal J).

14. ibid; CMS Law, "ANI Media Pvt Ltd V. OpenAI OPCO LLC, 24 July 2026" (AI & Copyright Case Tracker).

15. BTG Advaya, "ANI Media v OpenAI: The Delhi High Court Weighs In" (2026).

16. ibid.

17. Mandal (n 1); Copyright Act 1957, s 52.

18. "Protecting the Intangible: Rethinking Intellectual Property Strategies for AI Systems" (The IP Press, 2025).

19. Digital Personal Data Protection Act 2023, ss 11-12 (right to information/access).

20. "Behind Closed Datasets: The Trade Secret-DPDP Conflict in AI Training Data" (Record of Law, 2026).

21. ibid, citing the breach-of-confidence line of authority including Priya Puri (n 7).

22. "Intellectual Property and Personal Data in AI Datasets" (n 2).

23. Department for Promotion of Industry and Internal Trade, Working Paper on Generative AI and Copyright (referenced in n 2); Draft National IPR Policy 2016.

24. Streamlinefeed, "Delhi High Court Weighs OpenAI Copyright Appeal Following Landmark Fair Dealing Ruling" (2026).


[1]Ramendra Mandal, "Privacy, Copyright, and Trade Secret: The Growing Conflict Between Personal Data Rights and AI Training in India" (NASSCOM Community, 2026).

[2]"Intellectual Property and Personal Data in AI Datasets Under India's DPDP Act 2023" (2026) IPR Trends.

[3]Intepat, "Intellectual Property Law for AI in India 2026" (FAQ: "Does India have a trade-secret statute that protects AI training data and model weights?").

[4]Indian Contract Act 1872, s 27.

[5]Burlington Home Shopping Pvt Ltd v Rajnish Chibber 1995 (61) DLT 6 (Delhi HC).

[6]Diljeet Titus v Alfred A Adebare 2006 (32) PTC 609 (Delhi HC).

[7]American Express Bank Ltd v Ms Priya Puri 2006 (110) FLR 1061 (Delhi HC).

[8]Coco v A N Clark (Engineers) Ltd [1968] FSR 415, as applied in Indian breach-of-confidence jurisprudence.

[9]Mandal (n 1).

[10]"Intellectual Property and Personal Data in AI Datasets" (n 2), Abstract.

[11]Intepat (n 3).

[12]"ANI Media v OpenAI: The Delhi High Court Weighs In" (Mondaq India, 2026).

[13]ANI Media Pvt Ltd v OpenAI OpCo LLC, Delhi High Court, order dated 24 July 2026 (Bansal J).

[14]ibid; CMS Law, "ANI Media Pvt Ltd V. OpenAI OPCO LLC, 24 July 2026" (AI & Copyright Case Tracker).

[15]BTG Advaya, "ANI Media v OpenAI: The Delhi High Court Weighs In" (2026).

[16]ibid.

[17]Mandal (n 1); Copyright Act 1957, s 52.

[18]"Protecting the Intangible: Rethinking Intellectual Property Strategies for AI Systems" (The IP Press, 2025).

[19]Digital Personal Data Protection Act 2023, ss 11-12 (right to information/access).

[20]"Behind Closed Datasets: The Trade Secret-DPDP Conflict in AI Training Data" (Record of Law, 2026).

[21]ibid, citing the breach-of-confidence line of authority including Priya Puri (n 7).

[22]"Intellectual Property and Personal Data in AI Datasets" (n 2).

[23]Department for Promotion of Industry and Internal Trade, Working Paper on Generative AI and Copyright (referenced in n 2); Draft National IPR Policy 2016.

[24]Streamlinefeed, "Delhi High Court Weighs OpenAI Copyright Appeal Following Landmark Fair Dealing Ruling" (2026).