Table detection in heterogeneous documents

94Citations
Citations of this article
113Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Detecting tables in document images is important since not only do tables contain important information, but also most of the layout analysis methods fail in the presence of tables in the document image. Existing approaches for table detection mainly focus on detecting tables in single columns of text and do not work reliably on documents with varying layouts. This paper presents a practical algorithm for table detection that works with a high accuracy on documents with varying layouts (company reports, newspaper articles, magazine pages,.. ). An open source implementation of the algorithm is provided as part of the Tesseract OCR engine. Evaluation of the algorithm on document images from publicly available UNLV dataset shows competitive performance in comparison to the table detection module of a commercial OCR system. Copyright 2010 ACM.

Cite

CITATION STYLE

APA

Shafait, F., & Smith, R. (2010). Table detection in heterogeneous documents. In ACM International Conference Proceeding Series (pp. 65–72). Association for Computing Machinery (ACM). https://doi.org/10.1145/1815330.1815339

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free