Nautilus – An End-To-End METS/ALTO OCR Enhancement Pipeline

1Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

When a digital collection has been processed by OCR, the usability expecta-tions of patrons and researchers are high. While the former expect full text search to return all instances of terms in historical collections correctly, the lat-ter are more familiar with the impacts of OCR errors but would still like to apply big data analysis or machine-learning methods. All of these use cases depend on high quality textual transcriptions of the scans. This is why the National Library of Luxembourg (BnL) has developed a pipeline to improve OCR for existing digitised documents. Enhancing OCR in a digital library not only demands improved machine learning models, but also requires a coherent reprocessing strategy in order to apply them efficiently in production systems. The newly developed software tool, Nautilus, fulfils these requirements using METS/ALTO as a pivot format. The BnL has open-sourced it so that other libraries can re-use it on their own collections. This paper covers the creation of the ground truth, the details of the reprocessing pipeline, its production use on the entirety of the BnL collection, along with the estimated results. Based on a quality prediction measure, developed during the project, approximately 28 million additional text lines now exceed the quality threshold.

Cite

CITATION STYLE

APA

Maurer, Y., Schneider, P., & Marschall, R. (2023). Nautilus – An End-To-End METS/ALTO OCR Enhancement Pipeline. LIBER Quarterly, 33(1). https://doi.org/10.53377/lq.13330

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free