US government backs OpenAI's fair-use case against the NYT
The Justice Department told a Manhattan court that training language models on copyrighted text should count as fair use, breaking with the Copyright Office.
Symbolic image: a courtroom during a recess, a staffer wheeling a cart of case binders past the bench while screens cycle through a document comparison.
The US Department of Justice has sided with OpenAI and Microsoft in The New York Times' copyright suit, arguing that training language models on protected text falls under fair use.
At a glance
- The US Justice Department supports the fair-use defense of OpenAI and Microsoft in a Manhattan federal court.
- The New York Times sued in December 2023, seeking billions in damages and the destruction of GPT-4 and related models.
- Core argument: training copies are never released, and outputs usually lack substantial similarity to the originals.
- The department rejects the US Copyright Office report, saying it carries no binding authority.
The US Department of Justice has thrown its weight behind OpenAI and Microsoft in the copyright case brought by The New York Times, telling a federal court in Manhattan that training a large language model on protected text can qualify as fair use.
The argument in the filing
The government's reasoning turns on a distinction the newspaper's complaint treats as a single act. Copies made during training are never released to the public, the filing says, while what a finished model produces “often if not always” lacks substantial similarity to the works it learned from.
To describe the learning step, the department reaches for a literary comparison: Joan Didion studied Hemingway's sentences without displacing them. Holding a model liable for the same kind of absorption would “stifle the creativity copyright law is supposed to protect,” the filing argues. It also states plainly that “human beings create original works using LLMs.”
Breaking with the Copyright Office
What the filing rejects matters as much as what it endorses. A US Copyright Office report prepared under former Register Shira Perlmutter took a far more cautious view of training on protected works. The Justice Department answers that the report carries no binding authority, putting the executive branch at odds with its own copyright agency.
The filing also points to the Kadrey ruling, where a court had already addressed how a model's learning process should be treated.
What the Times is asking for
The newspaper sued in December 2023, alleging that OpenAI trained its systems on Times articles without permission. It seeks damages running into the billions, plus the destruction of GPT-4 and related models built on that text. Claims from other rights holders have since been consolidated with the case.
What is not settled
The sources available for this report do not establish the procedural form of the filing, the name of the presiding judge, or how OpenAI and The New York Times have responded to it. A government filing does not bind the court: it is a legal position, not a ruling.
FAQ
What does fair use mean for AI training?
Fair use is a limit on US copyright that can permit unlicensed use. The Justice Department argues model training qualifies because the training copies are never published and the resulting outputs usually lack substantial similarity to the originals.
How much is The New York Times seeking from OpenAI?
The newspaper is seeking damages in the billions, along with the destruction of GPT-4 and related models to the extent they were trained on its articles. The sources available for this report do not give an exact figure.
Does the court have to follow the US government's position?
No. The filing sets out the department's legal view and does not bind the Manhattan court. The case remains undecided.