A Review on Document Information Extraction Approaches

0Citations
Citations of this article
47Readers
Mendeley users who have this article in their library.

Abstract

Information extraction from documents has become great use of novel natural language processing areas. Most of the entity extraction methodologies are variant in a context such as medical area, financial area, also come even limited to the given language. It is better to have one generic approach applicable for any document type to extract entity information regardless of language, context, and structure. Also, another issue in such research is structural analysis while keeping the hierarchical, semantic, and heuristic features. Another problem identified is that usually, it requires a massive training corpus. Therefore, this research focus on mitigating such barriers. Several approaches have been identifying towards building document information extractors focusing on different disciplines. This research area involves natural language processing, semantic analysis, information extraction, and conceptual modelling. This paper presents a review of the information extraction mechanism to construct a generic framework for document extraction with aim of providing a solid base for upcoming research.

Cite

CITATION STYLE

APA

Silva, K., & Silva, T. (2021). A Review on Document Information Extraction Approaches. In International Conference Recent Advances in Natural Language Processing, RANLP (Vol. 2021-September, pp. 174–179). Incoma Ltd. https://doi.org/10.26615/issn.2603-2821.2021_024

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free