Judge Demands OpenAI to Release 20 Million Anonymized ChatGPT Chats in AI Copyright Dispute
On January 5, 2026, a federal judge in New York ordered OpenAI to supply 20 million anonymized user logs from ChatGPT. The decision, rendered by District Judge Sidney H. Stein, aligns with Magistrate Judge Ona T. Wang's earlier ruling, despite OpenAI's…
On January 5, 2026, a federal judge in New York ordered OpenAI to supply 20 million anonymized user logs from ChatGPT. The decision, rendered by District Judge Sidney H. Stein, aligns with Magistrate Judge Ona T. Wang's earlier ruling, despite OpenAI's concerns over user privacy.
The ruling pertains to a copyright lawsuit involving AI, where plaintiffs allege infringement through the reproduction of copyrighted content. OpenAI had proposed a narrower search of logs referencing the plaintiffs' works, arguing that providing the full dataset was overly burdensome and posed privacy risks. However, Judge Stein disagreed, citing the adequacy of existing privacy safeguards.
The legal proceedings began in July 2025, when several news organizations, including The New York Times Co. and Chicago Tribune Co. LLC, sought access to 120 million logs. Their objective was to investigate potential copyright infringements by ChatGPT.
Initially, OpenAI submitted a dataset of 20 million logs, which the plaintiffs initially accepted but subsequently deemed insufficient. Magistrate Judge Wang ruled in favor of the plaintiffs in November, and Judge Stein's affirmation enforces this under a protective order incorporating de-identification protocols.
On January 5, 2026, a federal judge in New York ordered OpenAI to supply 20 million anonymized user logs from ChatGPT.
OpenAI referenced a previous Second Circuit case involving SEC wiretap disclosures, arguing against the log handover. However, Judge Stein highlighted the distinction, noting that ChatGPT logs encompass voluntary user inputs and are owned by the company.
This decision facilitates pretrial discovery in the case In re OpenAI, Inc. Copyright Infringement Litigation (No. 1:25-md-03143), which consolidates multiple suits from various entities alleging unauthorized use of content to train AI models. The outcome of this case could set significant precedents for similar cases, especially concerning the balance between discovery proportionality and data privacy.
The ruling underscores the judiciary's readiness to mandate the release of substantial evidence, even when anonymized, to scrutinize AI training methodologies. This has implications for both content creators, who gain tools to contest potential copyright violations, and technology companies, who face increased examination of their data handling practices.
Based on reporting by Cyber Security News.
