Extracting mathematical components directly from PDF documents for mathematical expression recognition and retrieval

Botao Yu; Xuedong Tian; Wenjie Luo

Conference Proceedings

Extracting mathematical components directly from PDF documents for mathematical expression recognition and retrieval

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2014) 8795 170-179

DOI: 10.1007/978-3-319-11897-0_20

4Citations

8Readers

Get full text

Abstract

PDF document gains its popularity in information storage and exchange. With more and more documents, especially the scientific documents, available in PDF format, extracting mathematical expressions in PDF documents becomes an important issue in the field of mathematical expression recognition and retrieval. In this paper, we proposed a method of extracting mathematical components directly from PDF documents rather than cooperating indirectly with corresponding images converted from PDF files. Compared with traditional image-based method, the proposed method makes full use of the internal information of PDF documents such as font size, baseline, glyph bounding box and so on to extract the mathematical characters and their geometric information. The experimental result shows the method could meet the needs of the following processing of mathematical expressions such as formula structural analysis, reconstruction and retrieval, and has a higher efficiency than traditional image-based ways.

Author supplied keywords

Cite

CITATION STYLE

APA

Yu, B., Tian, X., & Luo, W. (2014). Extracting mathematical components directly from PDF documents for mathematical expression recognition and retrieval. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 8795, pp. 170–179). Springer Verlag. https://doi.org/10.1007/978-3-319-11897-0_20

Extracting mathematical components directly from PDF documents for mathematical expression recognition and retrieval

Abstract

Author supplied keywords

Cite

Register to see more suggestions