Microsoft and OpenAI executives admit AI tools substitute for news articles and books

Microsoft and OpenAI executives testified their AI products substitute for copyrighted news and books, with internal data showing an 83-93% click-through drop for publisher content.

Categorized in: AI News Legal
Published on: Sep 18, 2026
Microsoft and OpenAI executives admit AI tools substitute for news articles and books

Microsoft and OpenAI executives admitted in sworn testimony that their AI products serve as substitutes for copyrighted news articles and books, and acknowledged the threat their technology poses to publishers and authors, according to court documents unsealed Thursday in the Southern District of New York. The disclosures arrive as the multidistrict litigation heads toward rulings on which copyright claims will proceed to trial.

Microsoft CEO Satya Nadella testified that conversations with chatbots have "substituted" for publisher sites. OpenAI senior executive Nick Turley wrote that "[o]ur products are largely substitutive, period," the news publishers' partly unsealed motion for summary judgment shows. Jack Clark, OpenAI's former policy director who later co-founded Anthropic, wrote of books: "The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon." He added, "Our work in this area will make people unemployed."

The unsealing follows nearly two weeks after the parties filed redacted motions. The Trump administration has urged the court to rule that training AI on copyrighted works constitutes fair use.

Internal alarms over piracy and paywalled content

Nadella also testified that downloading pirated content is "absolutely" illegal and that paywalled content should be licensed for AI training. If he had known OpenAI scraped and trained on paywalled information, he would have "invoked [Microsoft's right to] require[] OpenAI to retrain its models," the filings state. Microsoft knew OpenAI used an illegal online library of pirated books as early as 2019, when Sam Altman and former vice president Dario Amodei presented an early chatbot version to Bill Gates.

OpenAI employees internally described the pirate library as "sketchy af" and said it "violates copyright restrictions," yet attempted to conceal the piracy. The company later deleted two training corpuses from the pirate library in 2022 due to legal concerns, in an effort known internally as "Project Clear." When staff told OpenAI President Greg Brockman about a hack to bypass the New York Times paywall for scraping, Brockman responded, "ah nice."

Microsoft Director of Applied Science Dr. Brent Hecht said copying millions of news articles without permission could be the "largest theft of labor in human history." He added that if OpenAI and Microsoft prevail on a fair use defense, it would "make a complete mockery of the idea of 'fair use.'" A Microsoft spokesperson said Hecht's comments "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views."

Content suppression and the 'doom loop'

Immediately after news publishers filed their lawsuits, OpenAI built a filter to suppress output of their content but did not suppress content from entities that had not sued, the unsealed motion said. Hecht called this an "accidental cover up" because it would result in "people who have a right over the content having less visibility into what was used for training."

This dynamic creates what the publishers describe as a "doom loop" that hurts the market for creative works and limits the production of additional training material. Microsoft's data shows an 83-93% drop in click-through rates for Times' and Daily News' content, and 51-94% for Ziff Davis' domains.

'Horse trading' of articles and books

The unsealed redactions also detail what Microsoft and OpenAI internally called "horse trading" deals, where they supplied copies of articles and books to each other and occasionally sold content between themselves. A Microsoft project code-named "Project Taxi" gave OpenAI a compilation of "billions" of webpages gathered to support the Bing search engine. Under "Project Mango," OpenAI paid Microsoft to develop and operate a crawler to copy webpage content.

OpenAI also distributed book copies to third-party contractors for a project where contractors read and summarized novels, and the company maintained a central library containing some books not used for AI training.

Steven Lieberman of Rothwell Figg Ernst & Manbeck PC, counsel for plaintiff New York Daily News, said: "Throughout this case Defendants insisted that these documents be treated as confidential so that the public could not see them. Well, now the cat is out of the bag. Finally, the world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior."

Why this matters for legal professionals

The unsealed testimony strips away the fair use defense's public framing and exposes internal admissions that AI outputs substitute for the very works the models were trained on. For IP litigators, the statements from Nadella, Turley, and Clark provide direct evidence on the fourth fair use factor - the effect on the market for the copyrighted work. The documents also reveal a pattern of internal knowledge about pirated training data and deliberate suppression of outputs from litigious publishers, facts that could influence discovery strategy in pending and future copyright cases. As the multidistrict litigation moves toward trial, practitioners in AI for Legal and AI for Executives & Strategy should watch how the court weighs these admissions against the administration's fair use position.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)