Internal documents tie Bodleian texts to OpenAI training data

single source· 1 articles · confidence: low · first seen 2026-09-26 11:00 UTC

What this means for you

Nothing to act on. The report gives no fee, volume or licence terms, so it is not a template for buying archive data. If you license digitised university collections for training, the terms Oxford and OpenAI agreed are the part worth finding — they are not here.

Oxford University has let OpenAI train its models on digitised historical texts from the Bodleian Library, according to internal documents seen by the Guardian. Those documents say Bodleian material digitised by OpenAI was used to "populate the OpenAI training set". University staff raised concerns about the reputational risk of partnering with the company behind ChatGPT. The Guardian's report, published on 26 September 2026, gives no fee, no volume of material and no licensing terms, and does not say whether Oxford or OpenAI has commented.

Key facts

  • ·Oxford University allowed OpenAI to train its AI models on digitised historical texts from the Bodleian Library, according to internal documents. source
  • ·The internal documents state that Bodleian material digitised by OpenAI was used to "populate the OpenAI training set". source
  • ·University staff voiced concerns about the reputational risk of partnering with the company behind ChatGPT. source
  • ·The Guardian published the report on 26 September 2026. source
  • ·No fee, volume of material or licensing terms are given in the report. source

What the sources say

  • The Guardian AI — Reports internal documents showing Bodleian texts went into OpenAI's training data, plus staff unease at the tie-up.

Sources

The original reporting. Follow these — they did the work.

← the wire