# Oxford says Bodleian digitisation also contributes OpenAI training data

> Oxford says its Bodleian digitisation pilot contributes training data to OpenAI. Public project records detail scanning and transcription, but not the material or models involved.

By BIG CHANGE Editorial

Published: 2026-09-27T01:10:31.965Z
Updated: 2026-09-27T01:10:31.965Z
Canonical: https://bigchange.ai/blog/oxford-bodleian-digitisation-openai-training-data

![An unidentified open bound volume rests in an archival cradle beneath a mounted document camera, with library shelves behind it.](https://bigchange.ai/api/media/file/oxford-bodleian-digitisation-hero-v1.png)
AI-generated conceptual illustration by BIG CHANGE.

Oxford presented its work with OpenAI as a way to scan public-domain library material and make it easier for researchers to find. [The Guardian reported on 26 September](https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt) that internal university documents describe the digitised material as helping to populate an OpenAI training set. An Oxford spokesperson told the newspaper that staff had been open about the project contributing training data. The public project description explains the scanning and transcription research in detail, but leaves that second use undefined.

## The big change: digitisation also serves AI training

- **What changed:** Oxford has acknowledged a training-data purpose alongside its public digitisation pilot, following the Guardian’s report about internal documents. The public record still does not identify a model trained on Bodleian material.
- **Why it matters:** Libraries can gain searchable digital collections while a technology partner gains material for AI development. Researchers and the public need to know which collections serve each purpose to assess the exchange.
- **What to watch:** Oxford says it retains rights to the scans and plans to publish the digitised material openly. A collection-level account of what was shared with OpenAI, and for which training use, would make the arrangement easier to assess.

## The public project is about scanning and transcription

[Oxford’s March 2025 announcement](https://www.bodleian.ox.ac.uk/about/media/oxford-and-openai-launch) described a pilot to digitise public-domain Bodleian material that had not been available online. It named 3,500 global dissertations dating from 1498 to 1884 among the intended collections and said the result would become searchable and accessible to students and researchers worldwide. It did not describe a transfer into an OpenAI training set.

The [Bodleian’s project page](https://www.bodleian.ox.ac.uk/node/4611321) says the pilot began in February 2025 with OpenAI funding. Its workstreams test scanning equipment and methods, examine how optical character recognition and handwriting text recognition extract text from scans, and explore AI-assisted metadata and collection search. These tasks can produce digital images, transcriptions and catalogue records. Testing whether a system can read a scanned page does not by itself establish that the page or its transcription entered a model-training dataset.

The Bodleian says it has captured about 125,000 images from its Global Dissertations collection and created another 429,000 files from handwritten catalogue cards. Its page describes the dissertation sample as European and American PhD theses from the nineteenth and twentieth centuries; the 2025 announcement identified 3,500 dissertations dating from 1498 to 1884. The public pages do not explain how those descriptions map onto the material scanned or transferred. Their counts describe digitisation output, not a published inventory of OpenAI’s training data.

## What the new reporting establishes

The Guardian says internal documents use the phrase “populate the OpenAI training set” for Bodleian material digitised by OpenAI. It also reports that 125,000 images of historical dissertations had been shared with OpenAI by June 2025, and says other texts scanned included 10,000 sixteenth-century broadside ballads. The report does not establish that every scanned ballad became training data. The newspaper says university meeting minutes obtained through a freedom of information request record staff concerns about reputational and environmental risks of the partnership. Those internal documents have not been made available in the sources we could inspect, so their precise wording, dates and scope beyond the Guardian’s account remain unverified here.

A university spokesperson told the Guardian the material being digitised was modest in scale, out of copyright and available to OpenAI on a non-exclusive basis. The spokesperson said the Bodleian kept rights to the scans and planned to publish them openly online within months. The spokesperson also rejected the suggestion that the machine-learning element had been hidden, telling the paper staff had been open that the project would contribute training data. OpenAI told the Guardian it was proud to help historical knowledge be reflected in AI models; its published [NextGenAI announcement](https://openai.com/index/introducing-nextgenai/) describes the Bodleian work as digitising rare texts and using its API to transcribe them.

Taken together, the report and Oxford’s response support saying that a training-data contribution is part of this arrangement. They do not establish which collection items were placed in which dataset, whether images or transcribed text were used, whether any particular released model was trained on them, or the terms under which OpenAI can reuse them. A training set can be assembled before a specific model is trained on it. Oxford’s public digitisation description and OpenAI’s 2025 announcement do not fill those gaps.

## Records still needed to map the training use

Oxford’s project page says the team planned open-access reports on each of its five workstreams. A report or collection inventory that names the material supplied to OpenAI and describes its permitted training uses would let scholars compare the public-access benefit with the data supplied to the company. Until then, the clearest account of that second use is the Guardian’s document-based reporting and the responses Oxford and OpenAI gave it.

## Sources & further reading

The [Guardian’s 26 September 2026 investigation](https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt) is the source for the internal-document claim, the reported sharing of dissertation images, the meeting-minute concerns and the Oxford and OpenAI responses. Its underlying internal documents were not available for our independent inspection.

[Oxford’s 4 March 2025 collaboration announcement](https://www.bodleian.ox.ac.uk/about/media/oxford-and-openai-launch) sets out the public-domain pilot and its intended researcher access. It does not specify an OpenAI model-training use.

The [Bodleian Digitisation Research project page](https://www.bodleian.ox.ac.uk/node/4611321) details the pilot’s workstreams and image counts. Its figures describe scanning output; the page does not identify a training dataset or a model trained on it.

[OpenAI’s 4 March 2025 NextGenAI announcement](https://openai.com/index/introducing-nextgenai/) describes digitisation and API transcription at the Bodleian. It does not identify training-set contents or a model trained on the scans.

## Sources

- [The Guardian: Oxford lets OpenAI train its AI models on Bodleian Library](https://www.theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt) — Original reporting, 26 September 2026, on internal documents and university meeting minutes; includes Oxford and OpenAI responses. The underlying internal documents were not available for our independent inspection, so the training-set wording and sharing details are attributed to the Guardian.
- [Oxford and OpenAI launch collaboration to advance research and education](https://www.bodleian.ox.ac.uk/about/media/oxford-and-openai-launch) — Oxford's 4 March 2025 announcement describes a public-domain Bodleian digitisation pilot, intended researcher access and an initial dissertation collection. It does not specify a training-data use.
- [Bodleian Digitisation Research project](https://www.bodleian.ox.ac.uk/node/4611321) — The Bodleian's public project page describes scanning, OCR/HTR, metadata and discovery workstreams and gives image counts. These are digitisation outputs, not a published inventory of OpenAI training material.
- [OpenAI: Introducing NextGenAI](https://openai.com/index/introducing-nextgenai/) — OpenAI's 4 March 2025 announcement describes Bodleian digitisation and transcription using its API. It does not identify a training set or a model trained on the material.
