In this paper, we present a relationship extraction based methodology for table structure recognition in PDF documents. The proposed deep learning-based method takes a bottom-up approach to table recognition in PDF documents. We outline the shortcomings of conventional approaches based on heuristics and machine learning-based top-down approaches. In this work, we explain how the task of table structure recognition can be modeled as a cell relationship extraction task and the importance of the bottom-up approach in recognizing the table cells. We use Multilayer Feedforward Neural Network for table structure recognition and compare the results of three feature sets. To gauge the performance of the proposed method, we prepared a training dataset using 250 tables in PDF documents, carefully selecting the table structures that are most commonly found in the documents. Our model achieves an overall accuracy of 97.95% and an F1-Score of 92.62% on the test dataset.
CITATION STYLE
Adiga, D., Bhat, S., Shah, M., & Vyeth, V. (2019). Table structure recognition based on cell relationship, a bottom-up approach. In International Conference Recent Advances in Natural Language Processing, RANLP (Vol. 2019-September, pp. 1–8). Incoma Ltd. https://doi.org/10.26615/978-954-452-056-4_001
Mendeley helps you to discover research relevant for your work.