NYT Asks a Judge to Sanction OpenAI Over Deleted ChatGPT Logs. The Deposition Names 78 Million Conversations.
The New York Times and more than a dozen other publishers filed a sanctions motion accusing OpenAI of destroying evidence and deleting billions of ChatGPT responses under a court preservation order.

News Breakdown · FiscEdge Academy
On Thursday, the New York Times, the Daily News, the Chicago Tribune, Ziff Davis, the Center for Investigative Reporting and more than a dozen other publishers filed a motion in Manhattan federal court asking a judge to sanction OpenAI. The ask: monetary penalties, and special jury instructions telling the eventual trial jury that OpenAI deleted billions of ChatGPT responses after being ordered to preserve them.
The filing leans on an April deposition of OpenAI data privacy engineer John Vincent "Vinnie" Monaco. According to the publishers, Monaco's testimony revealed that OpenAI had already built a database of roughly 78 million de-identified ChatGPT conversations, assembled before the Times even filed suit, and used internally to gauge how much of the model's output was infringing on other people's work.
The dollar figure that usually leads these stories is missing here, and that's the point. This isn't a funding round or a settlement number. It's a discovery fight, and the signal under the headline is about what companies are willing to say in court about their own systems, and what happens when that turns out not to be true.
The claim that was doing all the legal work
For two years, OpenAI's defense in the Times case rested partly on a technical argument: the company said it lacked the tooling to search its training data and chat logs for copyrighted journalism at scale. That claim mattered because it shaped what OpenAI was and wasn't required to produce in discovery.
The publishers now say that claim was false on its face. Monaco's deposition allegedly describes an internal system called "Project Giraffe," including a filter nicknamed "Bloom" built to detect and log when ChatGPT was regurgitating training content, stood up shortly after the lawsuit landed. New York Daily News attorney Steven Lieberman put it bluntly: OpenAI has been "making misrepresentations" about its search capabilities for two years. OpenAI spokesperson Drew Pusateri denied the allegations, saying the Times' case has weakened and that publishers are now trying to "invade the privacy" of unrelated users. That dispute will play out in court. What's already public is that a well-resourced AI lab told a federal judge it couldn't do something its own engineers were reportedly doing internally.
Why this is a founder problem, not just an OpenAI problem
Most SaaS founders will never face a discovery motion this size. But the underlying mechanics apply the moment your product logs user inputs, fine-tunes on customer data, or stores model outputs anywhere you don't fully control. Three things carry over directly:
What you say about your own systems is discoverable. If you tell a customer, a regulator, or a court that you "can't" search or segment your data a certain way, that statement needs to be true today, not true when you wrote your privacy policy. Engineering reality and legal representations drift apart fast when nobody's job is to keep them aligned.
Preservation orders are not a formality. Once litigation or a credible legal threat is on the table, routine data deletion and log rotation policies that felt like housekeeping become evidence-destruction risk. Founders who treat retention schedules as a pure cost-control lever, with no legal review trigger, are building the exact exposure OpenAI is now defending against.
Training data provenance is now a real liability line, not a footnote. Whether you're fine-tuning on scraped or licensed data, or just building on top of a foundation model whose own sourcing is under active litigation, the legal risk sits above you in the stack and can move downward. Knowing where your model's training data came from, and what your vendor has actually represented about it, belongs in the same due-diligence pass you'd run on a cap table, the kind of judgment call FiscEdge's AI for entrepreneurs track walks through before you commit to a foundation model.
None of this requires a legal department. It requires the same discipline FiscEdge teaches around building a legally sound business from day one: write down what your systems can and can't do, keep that document current, and treat data retention as a decision with legal weight, not just an infrastructure setting. If you're building the product itself on top of frontier models, the same logic extends into how you architect an AI-powered SaaS company that can survive a vendor's legal exposure becoming your own.
If you remember one thing
The most expensive sentence in this whole case might turn out to be "we don't have the capability to do that." Before you say it to a customer, a regulator, or a court, make sure it's still true, and make sure someone would still say it's true if a deposition happened tomorrow.
We cover the legal and operational fundamentals every founder needs in FiscEdge's Business Fundamentals course, and how to build defensibly on top of frontier AI models in Building SaaS with AI. Browse the full blog for more News Breakdowns. Follow @fiscedge for daily Business & AI analysis.
How interesting did you find this article?
The week's breakdowns, every Sunday.
Business & AI news decoded for founders. One email a week, no fluff.
Stay connected with FiscEdge Academy
Want more breakdowns like this one? Follow us and keep learning.