Treatment of text extracted from digital books for search engine indexing

1Citations
Citations of this article
1Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This article presents a methodology for treating texts extracted from digital books from Embrapa's 500 Questions 500 Answers Collection to index their content and to allow its access via a search engine. The methodology involves extracting the essential elements of the books, such as images and HTML files; pre-processing them; analyzing and editing them; and building suitable components for their indexing. In addition to a large amount of human analysis, the technologies used are Epub format for digital books, the Sigil editor, scripts for text processing, web representation standards, and Elasticsearch. The results show that this method can provide well-formatted texts for indexing and use in search engines, giving a rich user experience and enabling the construction of new digital solutions. Therefore, such a digital curation is essential for adding value to digital resources and meeting specific user needs.

Cite

CITATION STYLE

APA

Vaz, G. J., da Veiga, P. H. R. da C., Caldas, R. G., Vidal, W. C. L., de Assis, C. P., Correa, J. L., & Moura, M. F. (2023). Treatment of text extracted from digital books for search engine indexing. Revista Ibero-Americana de Ciencia Da Informacao, 16(2), 311–328. https://doi.org/10.26512/rici.v16.n2.2023.42740

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free