Open Access Content and AI: Questions of Use, Rights and Responsibility

Open Access Content and AI: Questions of Use, Rights and Responsibility

Open access has made scholarly research available to people who could not previously reach it, and we should take seriously the possibility that this same body of work could help AI systems produce more useful answers. A model that has encountered research from a broad range of disciplines, countries and languages has a better foundation than one drawing mainly on whatever happens to be abundant elsewhere on the web. AI tools that retrieve articles while answering a question can also help people find research outside their own field. These are possibilities, however, rather than guarantees of reliability, because the system still has to interpret the work accurately, distinguish a study’s findings from its limitations, let the reader see where an answer came from, and acknowledge and link to the creator of the work. The distinction between using an article to train a model and retrieving that article to answer a particular question matters here because retrieval creates a clearer opportunity to identify the source and check the answer against it, while the contribution of an individual work to a trained model is much harder to trace. Creative Commons has argued that attribution is more practicable in retrieval-based systems, although it remains inconsistent in AI outputs[1].

As more AI developers seek access to scholarly content, the word open can conceal important differences between collections and individual works. An article may be free for anyone to read without carrying a licence that permits every kind of copying or redistribution, and even within a collection of openly licensed articles or books, the terms can differ. PubMed Central makes this particularly clear: it provides free access to a large body of research but says that much of its content remains under traditional copyright restrictions, identifies a separate open access subset for broader reuse, and directs anyone downloading articles systematically to specific services rather than the main website. Its article datasets also distinguish licences that permit commercial use from those limited to non-commercial use. A developer collecting everything visible on a site could miss those distinctions, while an author who agreed to make a paper or book accessible to readers may have understandable questions about its inclusion in a commercial AI product.

These issues deserve care because neither the presence of a copyright notice nor an open licence provides a simple answer to every question about AI training. Creative Commons describes the application of copyright and its licences to training as complex, and cautions that choosing a more restrictive licence may not reliably prevent that use. At the same time, the terms attached to a work can matter greatly for other forms of copying, redistribution, and display. Therefore, we should be wary of a general claim that scraping open access content is either automatically permitted or automatically an infringement. We need to ask which material was collected, what the collector did with it, which rights and exceptions apply, and whether the AI system can reproduce or substitute for the original work. These questions become especially significant for books, where reproducing a substantial passage can affect the value of the work in a way that merely identifying or summarizing it may not.

Rights are only part of the picture, because obtaining content at scale also places demands on the services that host it. OAPEN[2] and the Directory of Open Access Books[3] reported that aggressive AI-driven scraping caused unprecedented server load and service disruptions during 2025, even as they continued to support broad access to scholarly books. This is a concrete reminder that a service has to pay for its servers, staff, maintenance, and capacity, regardless of whether readers pay to download a book. arXiv[4] similarly asks large-scale users to work through its designated bulk-access routes, explaining that it has limited server capacity and that indiscriminate downloading can impose a financial burden and interfere with access for human readers. Crossref, whose open metadata API receives around a billion requests a month, revised its rate limits after request volumes tripled over five years and periods of instability affected the service. Crossref’s figures[5] describe overall API demand, so they are not presented as a measure of AI scraping, but they do show why an open service must manage automated use if it is to remain dependable for everyone.

These examples also complicate the idea that a service should simply block AI crawlers. Researchers rely on automated systems to discover and analyze literature, while librarians, public-interest projects and smaller developers may have far fewer resources than large AI companies to negotiate access individually. Closing a collection to all machine use could protect the server while limiting some of the uses that open access was intended to encourage.

Using scholarly content responsibly means attending to the state of the record as well as to the text of the original publication. A more workable approach might begin with clear license information attached to each item, established routes for bulk access, realistic request limits, and communication between machine users and the organizations maintaining the collections. For an AI tool that retrieves sources for a user, it should also include links and enough bibliographic detail to let readers inspect the original article or book and, where relevant, see later corrections or retractions. PubMed Central’s downloadable datasets, for example, include information identifying retractions, corrections and expressions of concern.

Ultimately, AI systems benefit from the research that scholars and publishers have worked to make accessible, and readers benefit when those systems help them find and understand research. The institutions hosting open collections cannot be expected to carry unlimited traffic and operating costs without a discussion of how machine users contribute to their sustainability. Authors should be able to understand how their work is used and credited, and readers should be able to move from an AI-generated summary back to the research on which it rests. Open access has expanded the circulation of knowledge because people invested in making that circulation possible. As AI becomes another major user of these collections, its developers need to take part in sustaining the systems they depend on, including correct attributions, and sustainable use and support of open infrastructures.

CLOCKSS is a dark archive, and so the content we preserve is not available for human or machine use. A dark archive maintains preserved copies of scholarly works so that if a publication becomes unavailable from its original source, and a defined trigger event occurs, the content can be made available again. For CLOCKSS, the growing use of scholarly content by AI raises questions that extend beyond access and reuse to what happens to the original work over time. We serve a community of librarians and publishers whose views on AI training, licensing and commercial use vary, and so we are exploring these issues without assuming that there is one answer that will satisfy everyone. What connects this discussion to our work is the need to preserve the authoritative article or book, including the argument as it was published, the evidence on which it rests, and any changes made to its status later. An AI-generated account may help someone discover or understand a piece of research, but if the original publication is no longer available, the reader loses the ability to examine its methods and subtleties or to see whether it has subsequently been corrected or retracted.

This means that AI companies have a stake in the continued existence of dark archives and of a wider array of systems that can identify and point readers toward original sources.

As we consider how AI developers might use scholarly collections responsibly, it seems useful to separate several questions that are often discussed together: what the license permits, how content is collected, whether an AI answer identifies its sources, and who supports the infrastructure on which that use depends. Clearer license information agreed methods for bulk access and reasonable limits on automated requests could help address some of the immediate difficulties, while better links to original publications would make AI-generated accounts easier to verify. There is also room to ask whether organizations that derive considerable value from scholarly collections should contribute to the systems that host and preserve them. Librarians, publishers, authors and AI developers may reach different conclusions about particular forms of reuse, but all depend, in different ways, on the continued availability of the scholarly record.

Resources:
https://www.bloomsbury.com/us/discover/bloomsbury-academic/blog/featured/the-future-of-open-access-and-artificial-intelligence/

https://scholarlykitchen.sspnet.org/2026/03/31/what-ai-asks-of-open-access/

https://www.library.jhu.edu/news/2025/10/open-access-publishing-and-ai-considerations-for-authors/

[1] https://creativecommons.org/2026/07/24/attribution-in-the-age-of-ai/

[2] https://oapen.hypotheses.org/2217

[3] https://katinamagazine.org/content/article/resource-advisor/2026/what-makes-digital-library-trusted-open-infrastructure

[4] https://info.arxiv.org/help/bulk_data/index.html

[5] https://www.crossref.org/documentation/retrieve-metadata/text-and-data-mining/

 

Scroll to Top