The legal cases discussed in the first article post of this series illustrate the kinds of challenges archivists might encounter as AI becomes ubiquitous in many aspects of our lives. Yet even if archivists find it difficult to remain aware of every potentially relevant case, it can be useful to understand several key points of copyright law. This can allow archivists with many different specialties to make prudent judgments about certain materials even if they have little or no legal background.
Copyright Liability
One of the primary grounds that can be used to bring forth a copyright lawsuit is the principle of “substantial similarity.” Can an individual who inputs a prompt into AI software be considered liable for the product? And how can archivists be sure that digital items in their collections were not formulated using sources that enjoy copyright protection? If archivists are able, at the time of acquiring a collection, to gain a better understanding of how creators utilized artificial intelligence tools, they can feel sure that materials being prioritized for digitization will not raise any copyright concerns.
“Can an individual who inputs a prompt into AI software be considered liable for the product?”
Legal experts have acknowledged that a crucial counterargument in many AI debates relates to the improvement of AI tools. AI tools have the potential to bring about great advancements in many fields besides the library and archives profession. Some point out that without extensive digital records being provided to AI software companies, it will become impossible to make meaningful advances to these AI technologies. As it is archivists’ responsibility to facilitate access to the materials that are part of their collections, they should not necessarily want to prevent the materials they have worked so diligently to digitize from becoming part of what researchers using AI might discover.
This speaks to the need for archivists to create, update, and adhere to a collections development policy that considers potential donors’ expectations with respect to AI. Furthermore, these conversations might need to include other stakeholders in addition to the donor – this might be the case if the copyright of an item is jointly held or if the copyright of that item has passed from the original creator to family member(s).
DACS provides guidelines for archivists who need to indicate any number of restrictions related to items in their collection. It is possible that even this will not escape the transformative influence of AI technologies. In the case of born-digital materials and physical items that have been digitized, it might become necessary for archivists to include restriction notes related to the digital storage of recently accessioned materials. If a collection oversees several digital platforms and it is known that one contains training materials for any number of AI software, clarification with respect to use will become an even more fundamental part of acquisition negotiations.
Archivists and Recognizing Red Flags
Some might argue that smaller archives or special collections are less likely to acquire the kinds of AI-generated materials that would raise any copyright red flags. In these instances, archivists might be among those in the best position to ensure that any materials from their collection being used to “train” AI software are in the public domain. By working to establish close relationships with donors and keeping careful accession records, archivists can ensure that there is no ambiguity regarding how newly acquired materials are meant to be used.
Artwork and rare materials are often prioritized for digitization, and this is understandable given the positive attention this can bring to the collection. Still, archivists navigating the challenges of AI must be careful to gain a definitive understanding of how donors envision their creations being digitized and subsequently accessed. Consent to digitization is by no means the same as consent to the potential distribution of substantially similar AI-generated creations.
“Consent to digitization is by no means the same as consent to the potential distribution of substantially similar AI-generated creations.”
The National Archives and Records Administration (NARA) is implementing numerous pilot programs to address some of the most pressing AI-related questions facing the archival profession. These programs will address not only the issues that Archives employees must confront as they seek to provide access, but also the administrative issues that arise as archives such as NARA must serve a growing staff. These initiatives that will utilize artificial intelligence include efforts to more efficiently declassify materials and create descriptive metadata. Other plans will assess how artificial intelligence can allow Archives users to search the records created by archivists more easily.
Massive transcription efforts undertaken by the National Archives, the Library of Congress and the National Museum of American History make clear that these institutions recognize the importance of making their records available in a digital format. While this has often involved the extensive digitization of print records, the collections that archives receive and preserve going forward will necessarily include more born-digital material than ever before. Archivists who make increasing use of AI technologies will be responsible for finding new ways to encourage interest in their collections through other means, if AI is to complete the work previously done by volunteers giving generously of their time to support crowdsourcing efforts.
Personally Identifiable Information
A second area in which NARA has recognized artificial intelligence technologies to be of potential use is identifying PII (personally identifiable information) in born-digital records. Archivists completing records management work often need to determine which records contain this sensitive and confidential information and place access restrictions on parts of a collection. Yet in this dawning era of AI, ensuring that these records are kept out of the databases used to train AI software becomes more than a matter of copyright. It becomes a matter of maintaining the confidentiality of donors and the individuals described in the records those donors entrusted to them.
Because the National Archives is home to such a wide range of both print and digital collections, archives facing a range of challenges will surely be able to benefit in one way or another from the results of NARA’s pilot program. HAI (History Associates Incorporated) is also prepared to assist clients as they navigate the complexities of preserving collections that included items in physical and digital formats. HAI can offer the guidance of archivists whose expertise includes organizing born-digital collections and historians prepared to find answers to difficult questions of copyright and provenance.

Preserve and Activate Your Heritage
Organizations across sectors — from universities and museums to corporations, nonprofits, and government agencies — rely on strong heritage practices to protect the stories, collections, and records that shape identity and inspire future generations.
HAI's Heritage Services include:
- Archival Services & Collections Processing
- Digitization & Digital Preservation
- Records & Information Management
- Historical Research & Storytelling
- Exhibits, Interpretation & Digital Experiences
- ArchivalOne and PastView Partnership
- HeritageAdvantage CoLab
We help institutions safeguard their past so they can lead with clarity, credibility, and purpose. Let's preserve and activate your organization's heritage together.
Works Consulted:
Spoo, Robert. “Copyright Law and Archival Research.” Journal of Modern Literature 24, no. 2 (2000): 205–12. http://www.jstor.org/stable/3831907.
Dryden, Jean. “The Role of Copyright in Selection for Digitization.” The American Archivist 77, no. 1 (2014): 64–95. http://www.jstor.org/stable/43489586.
Hansen, David R. “Copyright Reform Principles for Libraries, Archives, and Other Memory Institutions.” Berkeley Technology Law Journal 29, no. 3 (2015): 1559–94. https://www.jstor.org/stable/26377576.
Fisher, Katherine. “Copyright and Preservation of Born-Digital Materials: Persistent Challenges and Selected Strategies.” The American Archivist 83, no. 2 (2020): 238–67. https://www.jstor.org/stable/48659901.
For the previous article in this series, click here.


