AI news ·
Judge allows authors' copyright infringement claims against Nvidia to proceed
A federal judge allowed a copyright lawsuit against Nvidia to proceed, rejecting most of the company's motion to dismiss. Three novelists claim Nvidia trained AI models on nearly 200,000 pirated books without permission.

Federal judge lets authors' copyright infringement case against Nvidia proceed
A federal judge in California has largely sided with three novelists who accused Nvidia of training its artificial intelligence systems on pirated books without permission. U.S. District Judge Jon Tigar denied most of Nvidia's motion to dismiss the proposed class action lawsuit on Tuesday, allowing claims for direct and contributory copyright infringement to move forward.
The authors - Brian Keene, Abdi Nazemian and Stewart O'Nan - filed suit more than two years ago. They say Nvidia trained multiple large language models using datasets that included their copyrighted works obtained illegally.
The dataset at the center of the case
The lawsuit focuses on a dataset called "The Pile," which contained a subcollection of nearly 200,000 pirated books called Books3. That collection came from Bibliotik, a shadow library hosting unauthorized copies of copyrighted works.
The authors claim Nvidia used The Pile to train several models in its Megatron line, including Megatron 345M, NeMo GPT-3 10B, and others. Books3 made up 12% of The Pile, and the authors' works appeared in Books3.
Nvidia argued that Megatron 345M was trained only on portions of The Pile that excluded Books3. The company submitted a screenshot from its website as evidence. Tigar rejected this approach at the pleadings stage, saying it could allow courts to dismiss valid claims before plaintiffs gather evidence through discovery.
Contributory infringement claim survives
The judge also allowed the contributory infringement claim to proceed. The authors alleged that Nvidia provided customers - including Writer, Persimmon AI Labs, and Amazon - with scripts designed to automatically download and preprocess The Pile for their own AI development.
Nvidia said the broader NeMo Megatron Framework had legitimate uses and the company never marketed it as a tool for copyright infringement. Tigar drew a distinction between the platform as a whole and the specific scripts in question.
"The scripts are alleged to have no other purpose than to speed up the process of infringement," he wrote.
The authors identified concrete instances of infringement by named customers rather than relying on speculation. Tigar found this sufficient to show Nvidia knew its tools were contributing to infringement by third parties.
One claim dismissed
The judge did dismiss the vicarious infringement claim. That claim requires showing a defendant had both the right to control infringing conduct and a direct financial interest in it.
The authors failed to explain how Nvidia could actually exercise control once a customer independently chose to access The Pile, Tigar found. They also did not establish that access to infringing material served as a draw for customers, rather than just an added benefit.
The judge gave plaintiffs 21 days to amend the dismissed claim.
What comes next
This case could reshape how generative AI and LLM companies acquire the massive datasets required to build their systems. The same law firm representing these authors also represents writers suing OpenAI over similar training data practices.
The decision suggests courts will scrutinize not just whether companies knew copyrighted material was in their training data, but also whether they provided tools specifically designed to facilitate infringement.