AI and web archiving workshop: AI4LAM affiliate event
Date: September 15, 2026
Location: Library of Congress
Time: 09:30-15:00
Artificial intelligence is rapidly reshaping how organizations collect, process, analyze, and provide access to web archives. This workshop brings together web archiving practitioners, researchers, and other interested community members to explore the current AI landscape in web archiving and discuss future opportunities for collaboration.
The event will combine short presentations and interactive discussions to map current projects, identify emerging use cases, and share lessons learned from ongoing experimentation. Participants will also have the opportunity to collectively explore future directions for AI in web archiving, including potential recommendations, best practices, community activities, and areas where further guidance may be needed.
Whether you are already applying AI in your work or are simply interested in understanding its potential impact on web archiving, we welcome your participation.
Program
Please note that this is a preliminary schedule and may change. We will distribute the final version once all submissions have been added.
09:30-10:00 | Welcome and introductions
- Workshop overview
- Key takeaways from the AI Literacy and Learning for LAMs Summit at the Library of Congress (July 2026)
10:00-11:00 | AI and web archiving: current landscape
- Overview of recent developments and literature
- Community presentations on:
- Current use cases
- Work in progress
- Future plans
- Participants interested in presenting are invited to submit a proposal via the workshop form.
11:00-11:15 | Break
11:15-12:00 | Visioning session
- What would we like AI to do for web archiving?
- Identifying opportunities, challenges, and unmet needs
12:00-13:00 | Lunch
13:00-14:30 | Collaborative workshop
- Interactive brainstorming and discussion sessions to explore:
- Priority areas for community action
- Potential recommendations and best practices
- Shared guidance, resources, and future activities
14:30-15:00 | Next steps and wrap-up
- Key findings and emerging themes
- Opportunities for continuing collaboration
Bibliography:
Authenticity
- Cargnelutti – Towards “deep fake” web archives? Trying to forge WARC files using ChatGPT (https://lil.law.harvard.edu/blog/2023/01/13/chatgpt-web-archives/)
Description
- Pandi, et al. – Can LLMs categorize the specialized documents from web archives in a better way? (https://www.cs.uic.edu/~cornelia/papers/2024.JCDL.pdf)
- Zhang – Towards a Digital Archivist: Applications of LLMs in Automated Web Archive Description (https://dl.acm.org/doi/10.1007/978-981-95-4861-3_40)
Discovery
- Balakireva – Evaluating Memento Service Optimizations (https://arxiv.org/abs/1906.00058)
- Deeds, et al. – GovScape: A Public Multimodal Search System for 70 Million Pages of Government PDFs (http://arxiv.org/abs/2511.11010)
- WARC-GPT
- Davis – Retrieval-Augmented Generation for Web Archives: A Comparative Study of WARC-GPT and a Custom Pipeline (https://journal.code4lib.org/articles/18555)
- Cargnelutti, et al. – WARC-GPT: An Open-Source Tool for Exploring Web Archives Using AI (https://lil.law.harvard.edu/blog/2024/02/12/warc-gpt-an-open-source-tool-for-exploring-web-archives-with-ai/)
Legal
- Taylor – AI for Temporal Web Forensics (https://nullhandle.org/blog/2025-08-20-ai-for-temporal-web-forensics.html)
Training Data
- Graham – Preserving The Web Is Not The Problem. Losing It Is. (https://www.techdirt.com/2026/02/17/preserving-the-web-is-not-the-problem-losing-it-is/)
- Nagel and Vaughan – Robots.txt and Crawler Politeness in the Age of Generative AI (https://digital.library.unt.edu/ark:/67531/metadc2472442/)
