The Authors Guild and other book plaintiffs are asking a federal judge to examine how OpenAI obtained copyrighted books separately from what it later did with them. Their September 17 public filings describe internal discussions about Library Genesis, the source of book collections used in model training, and argue that downloading, training and transferring copies each require their own copyright analysis.

OpenAI disputes that approach. In its own revised filing that day, it argues that acquiring books and training models served a single transformative purpose protected by fair use. The competing requests remain before Judge Sidney H. Stein in New York. A September 24 order extended the briefing schedule; it did not decide the copyright dispute. (Plaintiffs' brief, OpenAI's brief, scheduling order)

The big change

  • What changed: Authors can now point readers to a fuller public account of the book acquisition decisions they are challenging. The September filings expose more of the dispute over where training books came from and how companies handled them.
  • Why it matters: For writers seeking payment and developers defending their use of books, the legal treatment of acquisition matters alongside the treatment of training. The parties disagree about whether a model's purpose can justify earlier copying.
  • What to watch: The next briefs will contest both that legal argument and the evidence of harm to authors. Which evidence the judge permits, and which uses he evaluates separately, will shape the claims he can resolve without trial.

Books1 and Books2 in the public record

The September 17 documents are fuller public versions of summary-judgment submissions. The court's September 3 sealing order required parties to refile their briefs and factual statements with material unredacted where no party or third party sought continued sealing. Some passages remain blacked out.

The Authors Guild's September 21 release drew attention to two of those documents: the plaintiffs' legal brief, Docket 1982, and their corrected statement of facts, Docket 1987. The release is the Guild's account as a plaintiff.

In Docket 1987, the authors cite testimony and internal messages about naming training datasets Books1 and Books2. Paragraph 271 quotes former OpenAI employee Ben Mann identifying both as based on LibGen. Paragraph 275 quotes him explaining the wording in a draft paper: "it's deliberately vague since it’s libgen."

The same filing, at paragraph 749, reproduces a May 2020 message from then-OpenAI policy director Jack Clark anticipating competition with genre fiction authors. The plaintiffs present it as evidence of an internal concern about displacement. The message does not measure subsequent lost book sales or establish the court's view of market harm.

The document's title includes the words "undisputed material facts." That is the plaintiffs' submission for summary judgment. Under Local Civil Rule 56.1, the opposing party responds to the numbered statements with admissions or denials and supporting evidence. The title does not mean the judge has accepted the entire account.

The dispute over acquisition and training

The plaintiffs' brief asks the court to treat obtaining books without payment, copying them for training and distributing copies to other parties as distinct uses. It also argues that Microsoft bears responsibility for OpenAI's alleged infringement through its commercial relationship and ability to supervise the conduct.

On training itself, the authors argue that models exploit books' creative expression to produce competing work. They say that threatens authors' markets and deprives them of licensing revenue. (Plaintiffs, pages 25 to 38)

OpenAI's September 17 brief acknowledges using Library Genesis compilations to train GPT-3 and GPT-3.5. Its legal argument is that downloading was part of the process of building models that learn general language patterns. It says the court should assess the purpose of that whole process, and disputes both substitution for the books and legally recognizable market harm. (OpenAI, pages 3, 23 to 25 and 28 to 37)

Microsoft's revised book-case brief likewise defends training as fair use. It says Microsoft had no involvement in acquiring the LibGen datasets OpenAI used. On outputs, it cites an analysis of 8.2 million Copilot conversations that found 24 responses containing 30 matching words from the authors' asserted works. That is Microsoft's account of a specific analysis, with a particular matching threshold; it does not measure every way an author might claim harm.

The disagreement matters because US copyright law asks courts to consider the purpose of a use, the nature of the work, the amount taken and the effect on its market. A dispute about acquiring a copy and one about competition from generated text raise different factual questions. The parties are also arguing over how those questions fit together legally.

The evidence is still being contested

On September 23, OpenAI and Microsoft asked the court to exclude a supplemental expert report and portions of the authors' submissions relying on a study of AI-generated books and market dilution. Their motion's supporting brief alleges that plaintiffs' counsel funded the research and failed to disclose and present it through the required expert process. Those allegations are the defendants' argument for exclusion, not findings of misconduct. The motion directly targets material cited in Dockets 1982 and 1987.

Summary judgment allows a court to decide a claim or part of one when no genuine dispute over a material fact requires trial and the moving party is entitled to judgment under the law. (Federal Rule 56)

The public docket reviewed through September 25 records continuing briefing and expert-evidence motions. Stein's September 24 order sets October 23 for all parties' opposition briefs, October 30 for amicus briefs and November 20 for replies. For authors and AI developers, the next stage is the court's assessment of these competing accounts of copying and harm. The newly public filings supply arguments and cited evidence for that assessment; they supply no award of compensation or final fair-use ruling.