Digital weight watching: Reconstruction of scanned documents

Maarten Marx; Tim Gielissen

Journal ArticleOPEN ACCESS

Digital weight watching: Reconstruction of scanned documents

International Journal on Document Analysis and Recognition (2011) 14(2) 229-239

DOI: 10.1007/s10032-010-0135-3

3Citations

14Readers

Abstract

A web portal providing access to over 250. 000 scanned and OCRed cultural heritage documents is analyzed. The collection consists of the complete Dutch Hansard from 1917 to 1995. Each document consists of facsimile images of the original pages plus hidden OCRed text. The inclusion of images yields large file sizes of which less than 2% is the actual text. The search user interface of the portal provides poor ranking and not very informative document summaries (snippets). Thus, users are instrumental in weeding out non-relevant results. For that, they must assess the complete documents. This is a time-consuming and frustrating process because of long download and processing times of the large files. Instead of using the scanned images for relevance assessment, we propose to use a reconstruction of the original document from a purely semantic representation. Evaluation on the Dutch dataset shows that these reconstructions become two orders of magnitude smaller and still resemble the original to a high degree. In addition, they are easier to speed-read and evaluate for relevance, due to added hyperlinks and a presentation optimized for reading from a terminal. We describe the reconstruction process and evaluate the costs, the benefits, and the quality. © 2010 The Author(s).

Author supplied keywords

Cite

CITATION STYLE

APA

Marx, M., & Gielissen, T. (2011). Digital weight watching: Reconstruction of scanned documents. International Journal on Document Analysis and Recognition, 14(2), 229–239. https://doi.org/10.1007/s10032-010-0135-3

Digital weight watching: Reconstruction of scanned documents

Abstract

Author supplied keywords

Cite

Register to see more suggestions