Oxford presented its work with OpenAI as a way to scan public-domain library material and make it easier for researchers to find. The Guardian reported on 26 September that internal university documents describe the digitised material as helping to populate an OpenAI training set. An Oxford spokesperson told the newspaper that staff had been open about the project contributing training data. The public project description explains the scanning and transcription research in detail, but leaves that second use undefined.
The big change: digitisation also serves AI training
- What changed: Oxford has acknowledged a training-data purpose alongside its public digitisation pilot, following the Guardian’s report about internal documents. The public record still does not identify a model trained on Bodleian material.
- Why it matters: Libraries can gain searchable digital collections while a technology partner gains material for AI development. Researchers and the public need to know which collections serve each purpose to assess the exchange.
- What to watch: Oxford says it retains rights to the scans and plans to publish the digitised material openly. A collection-level account of what was shared with OpenAI, and for which training use, would make the arrangement easier to assess.
The public project is about scanning and transcription
Oxford’s March 2025 announcement described a pilot to digitise public-domain Bodleian material that had not been available online. It named 3,500 global dissertations dating from 1498 to 1884 among the intended collections and said the result would become searchable and accessible to students and researchers worldwide. It did not describe a transfer into an OpenAI training set.
The Bodleian’s project page says the pilot began in February 2025 with OpenAI funding. Its workstreams test scanning equipment and methods, examine how optical character recognition and handwriting text recognition extract text from scans, and explore AI-assisted metadata and collection search. These tasks can produce digital images, transcriptions and catalogue records. Testing whether a system can read a scanned page does not by itself establish that the page or its transcription entered a model-training dataset.
The Bodleian says it has captured about 125,000 images from its Global Dissertations collection and created another 429,000 files from handwritten catalogue cards. Its page describes the dissertation sample as European and American PhD theses from the nineteenth and twentieth centuries; the 2025 announcement identified 3,500 dissertations dating from 1498 to 1884. The public pages do not explain how those descriptions map onto the material scanned or transferred. Their counts describe digitisation output, not a published inventory of OpenAI’s training data.
What the new reporting establishes
The Guardian says internal documents use the phrase “populate the OpenAI training set” for Bodleian material digitised by OpenAI. It also reports that 125,000 images of historical dissertations had been shared with OpenAI by June 2025, and says other texts scanned included 10,000 sixteenth-century broadside ballads. The report does not establish that every scanned ballad became training data. The newspaper says university meeting minutes obtained through a freedom of information request record staff concerns about reputational and environmental risks of the partnership. Those internal documents have not been made available in the sources we could inspect, so their precise wording, dates and scope beyond the Guardian’s account remain unverified here.
A university spokesperson told the Guardian the material being digitised was modest in scale, out of copyright and available to OpenAI on a non-exclusive basis. The spokesperson said the Bodleian kept rights to the scans and planned to publish them openly online within months. The spokesperson also rejected the suggestion that the machine-learning element had been hidden, telling the paper staff had been open that the project would contribute training data. OpenAI told the Guardian it was proud to help historical knowledge be reflected in AI models; its published NextGenAI announcement describes the Bodleian work as digitising rare texts and using its API to transcribe them.
Taken together, the report and Oxford’s response support saying that a training-data contribution is part of this arrangement. They do not establish which collection items were placed in which dataset, whether images or transcribed text were used, whether any particular released model was trained on them, or the terms under which OpenAI can reuse them. A training set can be assembled before a specific model is trained on it. Oxford’s public digitisation description and OpenAI’s 2025 announcement do not fill those gaps.
Records still needed to map the training use
Oxford’s project page says the team planned open-access reports on each of its five workstreams. A report or collection inventory that names the material supplied to OpenAI and describes its permitted training uses would let scholars compare the public-access benefit with the data supplied to the company. Until then, the clearest account of that second use is the Guardian’s document-based reporting and the responses Oxford and OpenAI gave it.
Sources & further reading
The Guardian’s 26 September 2026 investigation is the source for the internal-document claim, the reported sharing of dissertation images, the meeting-minute concerns and the Oxford and OpenAI responses. Its underlying internal documents were not available for our independent inspection.
Oxford’s 4 March 2025 collaboration announcement sets out the public-domain pilot and its intended researcher access. It does not specify an OpenAI model-training use.
The Bodleian Digitisation Research project page details the pilot’s workstreams and image counts. Its figures describe scanning output; the page does not identify a training dataset or a model trained on it.
OpenAI’s 4 March 2025 NextGenAI announcement describes digitisation and API transcription at the Bodleian. It does not identify training-set contents or a model trained on the scans.



