The University of Oxford has granted OpenAI access to digitised historical texts from its Bodleian Library for AI model training, according to internal documents obtained via a freedom of information request. The arrangement, which began as a digitisation partnership announced in March 2025, was not publicly disclosed as a training data deal - a detail that has drawn concern from university staff over reputational risk and environmental commitments.
The Bodleian material has been used to "populate the OpenAI training set," the documents show. An OpenAI spokesperson said the company was "proud" to ensure "the AI models of today preserve the world's historical knowledge for the future," adding that with over a billion users, the technology should reflect "different cultures, histories and perspectives."
What Oxford shared with OpenAI
By June 2025, 125,000 images scanned from historical dissertations had been shared, including 19th and 20th-century PhD theses from European and American universities. Other digitised texts include a rare collection of 10,000 16th-century broadside ballads - song lyrics and musical notes once circulated on Tudor street corners. Staff have also discussed digitising 18th-century Irish state papers, the private letters of novelist Marie Edgeworth, and Dorothy Hodgkin's penicillin notebooks.
The Bodleian holds 23 million items, and meeting minutes reference the prospect of mass digitisation under the OpenAI contract, along with plans for an "Ask the Bod" chatbot. A university spokesperson said the amount of text being digitised was "modest in scale" and covered only out-of-copyright material. The Bodleian retains rights to the scans and will begin publishing them openly online within months.
Concerns inside the university
Minutes from Bodleian governance committee meetings record staff unease about partnering with OpenAI, specifically the reputational risk and the effect on Oxford's environmental commitments given the energy demands of AI infrastructure. The spokesperson rejected suggestions that the machine-learning element had been hidden, saying digitisation was the primary interest but that staff had been open about the project also contributing training data.
The hunt for offline data
Oxford is the only UK member of NextGenAI, a project through which OpenAI has struck similar agreements with the Boston Public Library, Caltech, MIT, and the University of Michigan. The deals reflect a broader industry shift: scraped websites are increasingly saturated with AI-generated content, making them less useful for training. Developers have turned to physical book collections - often historical, often obscure.
Secondhand booksellers have reported a surge in orders for titles unlikely to exist in digitised form, such as a guide to 18th-century African agricultural implements or biographies of 1950s car drivers. Rival AI company Anthropic has spent tens of millions of dollars acquiring books, slicing off their spines for scanning, and then pulping them. Anthropic has said it does not buy and destroy rare and antiquarian books. A separate investigation by 404 Media used a tracking device to trace a secondhand book order to an Amazon facility where books were dismantled and scanned.
Why this matters for education, legal, and writing professionals
For educators and researchers, the Oxford deal signals that institutional library collections are becoming training fodder for commercial AI systems - often without clear public disclosure. The Bodleian's arrangement keeps the physical collections intact, unlike some industry practices, but the reputational and copyright questions remain unresolved. Legal professionals should watch how "out of copyright" status interacts with data-rights claims when digitised works are used for machine learning. Writers and publishers face a marketplace where obscure print works are suddenly valuable as fresh training data, raising questions about consent, attribution, and the downstream use of digitised cultural heritage.
Professionals who work with AI tools or manage institutional data may want to understand how training datasets are sourced. For those building skills in this area, Generative AI Courses cover how large language models are trained and deployed. Research staff navigating these partnerships can explore AI Research Assistant Courses for practical guidance on working with AI systems in academic settings.
Your membership also unlocks: