A summary judgment brief unsealed on 17 September 2026 in the copyright case that news organizations brought against OpenAI and Microsoft quotes internal documents in which a Microsoft executive described scraping news for AI training as "an astonishing theft of unprecedented proportions", and in which OpenAI's head of ChatGPT said publishers face an "existential threat".
The document is the news plaintiffs' combined summary judgment brief in In re OpenAI, Inc., Copyright Infringement Litigation, MDL No. 1:25-md-3143 (SHS) (OTW), filed in the US District Court for the Southern District of New York on 17 September and running to 92 pages. The plaintiffs include The New York Times, the New York Daily News titles, the Center for Investigative Reporting and Ziff Davis. They ask the court to rule before trial that the two companies infringed their copyrights at three stages: acquiring articles, training models on them, and, for Microsoft, using retrieved copies to ground chatbot answers.
What the brief puts on the record
Microsoft's director of applied science, Dr. Brent Hecht, wrote in a January 2023 memo that scraping news for training was "an astonishing theft of unprecedented proportions" and perhaps "the largest theft of labor in human history", according to the filing. He also said, the plaintiffs argue, that the practice would "make a complete mockery of the idea of fair use" - the doctrine on which both companies base their defense of unlicensed training.
OpenAI's head of ChatGPT, Nick Turley, is quoted as writing that publishers face an "existential threat" from products that are "largely substitutive, period" and that "will get more and more substitutive as they get better". A Microsoft internal document went further, describing a "doom loop": "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain."
The brief also cites OpenAI co-founder and president Greg Brockman, who wrote that the models were "excellent at news" and replied "ah nice" when a researcher, Nick Ryder, told him about what he called a hack to get around the Times' paywall. Microsoft chief executive Satya Nadella said in a deposition that "anything that is paywalled should be licensed" for grounding or training, and agreed that chatbots have substituted for users going to the underlying source.
The traffic baseline and the dataset counts
The click-through figures come from Microsoft's own data, comparing its Copilot "answer engine" with traditional Bing search: a decline of 87 to 93 percent for The Times, 83 to 91 percent for the Daily News titles and 51 to 94 percent for ZD.
On scale, the filing states that OpenAI's mid-training datasets contain more than 91,692 copies of the plaintiffs' works; that a dataset assembled under a Microsoft-OpenAI program named Project Mango contains copies of at least 160,903 unique works; and that a New York Times Annotated Corpus obtained from a third party held more than 1.8 million articles published between January 1987 and June 2007, under a license limited to non-commercial research. One post-filtered Common Crawl set used for training contained 2,064,805 items from nytimes.com, the brief says.
Why it matters
Fair use turns in part on whether the copying substitutes for the original work. The plaintiffs argue that the quotes and the traffic data show substitution, and that licensing markets the companies could have used instead were superseded; they are seeking statutory damages for each infringed article. If the court finds liability, the question of what AI developers must pay for news content moves from negotiation to legal risk for every model trained on publisher data.
What the filing does not establish
Two limits are worth keeping in view. First, the descriptions above come from the plaintiffs' brief, which quotes the internal documents as excerpts rather than in full, and both companies dispute how they should be read. Second, no ruling has been issued, and the motion is deliberately narrow: the plaintiffs seek judgment only on works whose outputs show extensive verbatim overlap, leaving their remaining claims to trial.
Microsoft pushed back, saying the documents authored by Hecht "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views", and that its products are a transformative fair use that does not substitute for news sites. OpenAI did not immediately respond to Ars Technica's request for comment. Steven Lieberman, counsel for the New York Daily News and seven sister papers, said the newly visible material shows that "OpenAI and Microsoft knew that what they were doing was wrong".
Sources