Transparent Practices: OCR and AI in the Archives

  • Hastings R
  • Weymouth A
N/ACitations
Citations of this article
1Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This paper examines optical character recognition (OCR) through the lens of archival ethics as outlined in the Society of American Archivists (SAA) Core Values Statement and Code of Ethics, given the current debates surrounding artificial intelligence (AI). A literature review highlights persistent challenges of authenticity and integrity, transparency and accountability, access and equity, and responsible stewardship and sustainability, as well as new concerns about bias, sustainability, and accountability using large language models (LLM). A case study describes systematic testing of LLM, transformer model (TM), and neural network (NN) architectures and examines the challenges in creating a reliable, scalable in-house OCR tool named Opticolumn. This case study finds that NN approaches better align with archival ethics than do LLM tools, which may generate fabrications, but that OCR tool choice will depend on the capacities and preferences of individual institutions.

Cite

CITATION STYLE

APA

Hastings, R., & Weymouth, A. (2026). Transparent Practices: OCR and AI in the Archives. Collections: A Journal for Museum and Archives Professionals, 22(2), 130–152. https://doi.org/10.1177/15501906261439241

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free