Abstract
This article presents a methodology for treating texts extracted from digital books from Embrapa's 500 Questions 500 Answers Collection to index their content and to allow its access via a search engine. The methodology involves extracting the essential elements of the books, such as images and HTML files; pre-processing them; analyzing and editing them; and building suitable components for their indexing. In addition to a large amount of human analysis, the technologies used are Epub format for digital books, the Sigil editor, scripts for text processing, web representation standards, and Elasticsearch. The results show that this method can provide well-formatted texts for indexing and use in search engines, giving a rich user experience and enabling the construction of new digital solutions. Therefore, such a digital curation is essential for adding value to digital resources and meeting specific user needs.
Author supplied keywords
Cite
CITATION STYLE
Vaz, G. J., da Veiga, P. H. R. da C., Caldas, R. G., Vidal, W. C. L., de Assis, C. P., Correa, J. L., & Moura, M. F. (2023). Treatment of text extracted from digital books for search engine indexing. Revista Ibero-Americana de Ciencia Da Informacao, 16(2), 311–328. https://doi.org/10.26512/rici.v16.n2.2023.42740
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.