By HAI Staff
Archivists have not been able to escape the question that we all seem to be confronted with in 2025: how is AI changing the work that you do? In a profession such as ours that so often requires us to handle materials directly, artificial intelligence can seem full of potential rather than a direct threat to the work we do to preserve the past. Yet it is far from the case that using AI to create and fill digital archives will not complicate our work. In particular, many archivists might remain uncertain when confronted with newly raised questions surrounding copyright.
“[A]rtificial intelligence can seem full of potential rather than a direct threat to the work we do to preserve the past.”
Copyright Meets AI
Archivists and librarians, especially those responsible for overseeing the acquisition of new collections, have long had to keep aware of ever-changing U.S. copyright law. Archivists responsible for complying with records management guidelines must remain especially attuned to these changes. This is especially true as decisions related to the weeding and digitization of collection materials must be made as these enter the public domain. Copyright-related questions arise as archivists determine which records they can make available in a digital format: when artificial intelligence tools factor into the creation of these records, the questions are only likely to multiply.
One key element that archivists typically use while describing collections items is the creator of those item. Artificial intelligence has complicated the very nature of what used to be simply another required DACS element. A person who produces a document using a tool such as ChatGPT might in a sense be considered the “creator” of that item. However, for an archivist to identify that person as the creator of that document would by no means be painting the whole picture. In another sense, the true creator(s) of this document might be considered the creators of the documents supplied to train the AI software and upon which that software relies. The striking similarities between certain AI-generated articles and past publications have led to recent lawsuits that can offer archivists a clue as to the kind of challenges they might encounter in the future.
Digital Dilemma
Since archivists are increasingly required to work with born-digital records, it is almost certain that more and such records will be generated by artificial intelligence tools such as ChatGPT. The Copyright Office has already expressed its intent to offer copyright protections to works produced partially but not entirely produced by AI software. An increasing number of legal cases reflect the ever-changing views surrounding these new technologies.
A first case of note involved Hathi Trust, a vast digital repository utilized by many universities to make records from their archives and library collections accessible. In 2014, plaintiffs brought suit against HathiTrust for including searchable versions of some of their works in the database, unconvinced that the doctrine of fair use could be applied in this case. While AI was not the source of any of the works considered as part of this case, it serves as a caution to archivists today that claiming fair use doctrine can not always be applied in the case of our latest digital acquisitions.
Another recent case centered around a key American news source, The New York Times. The Times brought suit against OpenAI and Microsoft, citing these companies’ “unpermitted use of Times articles to train GPT large language models.” Representatives for Open AI continue to rely in part on questions of fair use in order to support their arguments; however, U.S. District Court Judge Sidney Stein has been inclined to side with New York Times representatives alleging infringement.
Archivists regularly encounter collections that include audiovisual material in addition to physical documents. Those who have studied copyright understand that these materials are covered by distinct kinds of copyright protections. The lawsuit brought by illustrators Sarah Andersen, Kelly McKernan and Karla Ortiz against Stability AI and other companies remind archivists that AI-generated images in their collections should be examined in light of many similar questions.
“Even if archivists take careful steps to ensure that there are no blurred lines with respect to the copyright holder of newly acquired materials, this becomes more challenging if AI-generated documents or files are included as part of digital acquisitions.”
Archivists are often responsible for overseeing the upload of materials into public-facing databases. For those acting in this role, cases related to the interpretation of the Digital Millenium Copyright Act (DMCA) should also be considered. Importantly, this Act outlines the kind of liability that platforms hosting copyright material might face if any claims are brought forward surround those materials. Even if archivists take careful steps to ensure that there are no blurred lines with respect to the copyright holder of newly acquired materials, this becomes more challenging if AI-generated documents or files are included as part of digital acquisitions. This serves as a reminder that not only do archives and libraries have a responsibility to ensure that the platforms they use meet these requirements: they can also play an active role in encouraging researchers and other users to take care to access digital records in a way that ensures compliance with copyright restrictions.
Ongoing Evolution
Continued refinement of the doctrine of fair use will further impact the way in which archivists go about their work. Past copyright cases exploring the implications of fair use recognize that educators (including archivists) are obligated to support the meaningful research being undertaken by those most likely to make use of that archive’s resources. Yet the reality of AI-generated documents constitutes an emerging challenge. Donors of a collection might well be sympathetic to the use of materials from that collection in an educational context. Still, items from a collection are digitized and made available in an archive’s digital repository, it becomes increasingly difficult to know how those materials will be used or in what context profoundly similar AI-generated materials will “appear.” Nor are such materials the only ones that merit archivists’ attention. The U.S. Copyright Office has devoted time to better understanding how AI technologies can create images and videos falsely attributed to or associated with a person. Archivists responsible for managing any number of collections containing an individual’s personal papers must take care to ensure that the physical and digital records they receive together create an accurate representation of that person’s life.
HAI’s CEO Beth Maser sat down and chatted with our Director of Archives and Information Management, Megan O’Hern-Crook in an article for Inc. Her conversation with Megan centers on AI’s dependence on data quality. Read about it here.

Preserve and Activate Your Heritage
Organizations across sectors — from universities and museums to corporations, nonprofits, and government agencies — rely on strong heritage practices to protect the stories, collections, and records that shape identity and inspire future generations.
HAI's Heritage Services include:
- Archival Services & Collections Processing
- Digitization & Digital Preservation
- Records & Information Management
- Historical Research & Storytelling
- Exhibits, Interpretation & Digital Experiences
- ArchivalOne and PastView Partnership
- HeritageAdvantage CoLab
We help institutions safeguard their past so they can lead with clarity, credibility, and purpose. Let's preserve and activate your organization's heritage together.
Cover Image: @cottonbro studios via Pexels


