For almost three years, the most interesting parts of The New York Times’ copyright case against OpenAI and Microsoft sat behind black bars. This week the bars came off. What emerged is the kind of document set that makes corporate counsel reach for the antacids: internal memos in which employees of the companies being sued describe the conduct at issue in language considerably harsher than anything the plaintiffs’ lawyers came up with on their own.
The Times sued in December 2023 in the Southern District of New York. The case, now before Judge Sidney Stein and consolidated with claims from the New York Daily News and the Center for Investigative Reporting, survived a motion to dismiss in early 2025 and has ground through discovery ever since. Both sides filed for summary judgment in early September. On September 17 and 18, mostly unredacted versions of those briefs became public. They are the first real look at what three years of depositions and document production actually turned up.
The number that should worry every publisher on the open web
Forget the quotable outrage for a second. The most legally dangerous item in the unsealed material is a percentage.
The plaintiffs’ brief cites a January 2024 internal Microsoft presentation finding that Copilot answers cut click-throughs to nytimes.com by as much as 93 percent compared with a traditional Bing search result. Not an estimate from a publisher trade group. Not a survey. Microsoft’s own telemetry, measuring Microsoft’s own product, arriving at a number that would be an extinction-level figure for any business that depends on referral traffic.
The presentation reportedly described the dynamic as a “doom loop”: the AI layer answers the question, the user never clicks, the publisher loses the revenue that funded the reporting, and the supply of reporting the AI depends on starts to thin out.
Why does this matter more than the juicy quotes? Because fair use turns on four factors, and the fourth, the effect on the market for the original work, is where most AI training cases will actually be decided. Defendants across this wave of litigation have argued that market harm is speculative. It is much harder to call a harm speculative when the defendant measured it, named it and put it in a slide deck.
A paper trail with names on it
What Microsoft employees wrote
The headline quote belongs to Brent Hecht, a Microsoft director of applied science, who in a January 2023 memo called the mass ingestion of human work “an astonishing theft of unprecedented proportions” and possibly “the largest theft of labor in human history.” That was written roughly a year before the presentation about collapsing click-throughs, and by the same person.
Then there is the chief executive. In deposition testimony cited in the brief, Satya Nadella said that anything paywalled should be licensed by anyone who wants to use it for grounding or training, and acknowledged that conversational AI has substituted for visits to original sources. Coming from the CEO of one of the two defendants, that is not a comfortable sentence to have read back to a jury.
What OpenAI employees wrote
OpenAI’s side of the ledger is no gentler. Nick Turley, who runs ChatGPT, is quoted describing an existential threat to publishers and writing that the company’s products are largely substitutive. Greg Brockman, OpenAI’s president, is quoted calling the models excellent at news. In a separate exchange over a method of getting around the Times’ paywall, raised by researcher Nick Ryder, Brockman’s recorded reply was two words: “ah nice.”
Two words. Prosecutors have built cases on less, and plaintiffs’ lawyers certainly have.
How much Times journalism is actually in there
The filings also put figures on the scale of ingestion. These are the plaintiffs’ numbers, drawn from discovery, and they describe what sat inside training corpora rather than what any model reproduced:
- More than 91,692 copies of works from the Times, the Daily News and the Center for Investigative Reporting in OpenAI’s mid-training datasets.
- More than two million documents from nytimes.com in a dataset derived from Common Crawl, the public web archive that underpins a great deal of modern model training.
- At least 160,903 unique news works in a dataset referred to internally as Project Mango.
- Two data-sharing efforts, code-named Project Taxi and Project Mango, through which material moved between the two companies.
The plaintiffs further allege that copyright management information was stripped from training material and that paywalled content was pulled via the Bing index. Worth keeping straight: a count of documents in a corpus proves ingestion, not infringing output. Those are different legal questions, and the second one is where the defense has its firmest ground.
The defense is stronger than the headlines suggest
Nobody has been found liable of anything.
Microsoft’s response to the Hecht material is that it reflects one employee’s individual perspective and is not a legal analysis or a statement of company position. That is a real argument. Large companies employ people paid specifically to write alarming internal memos, and an internal critic’s framing is not an admission.
On the merits, both companies maintain that training is transformative fair use and that their products do not substitute for journalism. The defense briefs point to large-scale analysis of Copilot logs indicating that very few responses substantially matched plaintiffs’ material, and argue that a product can reshape a market without violating anyone’s copyright.
There is also a thumb on the scale from Washington. In early September the Justice Department filed a statement of interest broadly supporting the position that training large language models on copyrighted work qualifies as fair use, framing it in terms of scientific progress and competition with China. It is the federal government’s first real intervention here, and judges will notice.
What this means
The unsealed briefs have not settled the fair use question. They have changed what the argument is about.
Until now, the AI industry’s strongest rhetorical position was that harm to publishers was hypothetical and that nobody could really prove otherwise. The filings make that harder to say with a straight face, because the evidence of substitution is not coming from the plaintiffs’ economists. It is coming from the defendants’ own product telemetry, their own executives and their own internal warnings, written well before the lawsuit.
That does not mean the Times wins. Courts have so far been reasonably receptive to transformative-use arguments about model training, and the gap between “we ingested your archive” and “we reproduced your archive” is a real one that OpenAI and Microsoft will keep pressing. Judge Stein could split the difference, finding training lawful while treating retrieval-based products that answer questions using fresh article content very differently.
The practical effect may arrive before any ruling does. Every publisher negotiating an AI licensing deal now has a specific number to put on the table, and it came from Microsoft. Every media executive who was told traffic loss could not be attributed to AI answers has a document suggesting otherwise. Whatever Judge Stein decides, the price of news as an input to AI systems just went up, because the people buying it have been shown to know exactly what it is worth.




