Extracting mathematical components directly from PDF documents for mathematical expression recognition and retrieval

4Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

PDF document gains its popularity in information storage and exchange. With more and more documents, especially the scientific documents, available in PDF format, extracting mathematical expressions in PDF documents becomes an important issue in the field of mathematical expression recognition and retrieval. In this paper, we proposed a method of extracting mathematical components directly from PDF documents rather than cooperating indirectly with corresponding images converted from PDF files. Compared with traditional image-based method, the proposed method makes full use of the internal information of PDF documents such as font size, baseline, glyph bounding box and so on to extract the mathematical characters and their geometric information. The experimental result shows the method could meet the needs of the following processing of mathematical expressions such as formula structural analysis, reconstruction and retrieval, and has a higher efficiency than traditional image-based ways.

Cite

CITATION STYLE

APA

Yu, B., Tian, X., & Luo, W. (2014). Extracting mathematical components directly from PDF documents for mathematical expression recognition and retrieval. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 8795, pp. 170–179). Springer Verlag. https://doi.org/10.1007/978-3-319-11897-0_20

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free