Newly unsealed court documents in Authors Guild v. OpenAI reveal that executives at OpenAI and Microsoft knowingly trained AI models on pirated books from "sketchy" sources and discussed the technology as a direct threat to human authors' livelihoods. The filings, made public last Thursday, include internal communications showing leaders understood their actions were illegal and would cause widespread job loss among writers - and proceeded anyway.
The class-action lawsuit, which includes authors George R.R. Martin, John Grisham, Jodi Picoult, and Jonathan Franzen alongside the Authors Guild, alleges the companies used copyrighted books without permission to train GPT models. The new briefs put internal Slack messages, emails, and deposition testimony at the center of the case, documenting what plaintiffs describe as a "lawless determination" to win the AI race regardless of legal or human cost.
Executives acknowledged the technology would replace writers
OpenAI Policy Director Jack Clark wrote in May 2020 that the company's work would "increasingly lead to us creating systems that substitute for the labor of people." He went further: "Our work in this area will make people unemployed. . . There will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns and release anyway."
In 2022, OpenAI hired Tarun Gogineni to improve its models' writing quality. Gogineni described his "research mission" as having GPT write the final two books of George R.R. Martin's A Song of Ice and Fire series, musing he would "rest easy knowing that even if GRRM dies early, GPT-5 will autocomplete his series." When authors complained their work was stolen and they were losing income to AI-generated competition, Gogineni called the complaints "acceptable economic disruption" and predicted "the death of the reader" as "machines create slop for more machines."
Microsoft knew about pirated training data from the start
The filings show Microsoft was aware OpenAI used LibGen - a repository of pirated books and academic papers - as early as April 2019. Sam Altman and Dario Amodei disclosed the use of LibGen to Bill Gates and Microsoft CTO Kevin Scott during a presentation of an early GPT-3 version, according to the documents.
When OpenAI researcher Sam McCandlish raised concerns, he framed the issue in terms of public relations rather than legality: "I was just worried about optics - i.e. 'openai uses copyrighted data from sketchy russian website' showing up on Hacker News would be unfortunate." Amodei, then Research Director, responded that LibGen was "a bit sketchier" as a training set.
Project Clear: the effort to delete evidence
In summer 2022, OpenAI launched what it internally called "Project Clear" to remove LibGen files from its systems. On June 15, an employee asked in a Slack channel: "general q: how concerned are we about mentions of libgen? (they're all over google docs/slack/github . . .)." That evening, VP of Research Bob McGrew responded: "Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage. What would be involved in that?"
Authors Guild CEO Mary Rasenberger said the filings "reveal shocking disdain for writers and their work through repeated, intentional decisions to steal books rather than pay for them with full knowledge that their products will destroy the careers of authors." The case is part of multidistrict litigation in Manhattan, with additional briefing expected in the coming months and a hearing anticipated in early 2027.
Why this matters for creatives and communications professionals
The internal communications documented in these filings are not abstract policy debates - they are direct evidence that companies building generative AI understood the harm to working writers and chose to proceed. For authors, journalists, and anyone whose income depends on intellectual property, this case will set precedent on whether training AI on copyrighted work without permission constitutes infringement. The outcome will shape licensing markets, fair use boundaries, and the economic viability of creative professions for years. Professionals in legal, PR, and policy roles should track this litigation closely, as it intersects with broader regulatory shifts covered in AI Public Policy Courses and AI Regulatory Compliance Courses.
Your membership also unlocks: