Abstract
Synonyms Information extraction; Text analytics Definition Information extraction (IE) is the process of au-tomatically extracting structured pieces of in-formation from unstructured or semi-structured text documents. Classical problems in informa-tion extraction include named-entity recognition (identifying mentions of persons, places, organi-zations, etc.) and relationship extraction (iden-tifying mentions of relationships between such named entities). Web information extraction is the application of IE techniques to process the vast amounts of unstructured content on the Web. Due to the nature of the content on the Web, in addition to named-entity and relationship extrac-tion, there is growing interest in more complex tasks such as extraction of reviews, opinions, and sentiments. Historical Background Historically, information extraction was studied by the Natural Language Processing community in the context of identifying organizations, lo-cations, and person names in news articles and military reports [15]. From early on, information extraction systems were based on the knowl-edge engineering approach of developing care-fully crafted sets of rules for each task. These systems view the text as an input sequence of symbols, and extraction rules are specified as regular expressions over the lexical features of these symbols. The formalism underlying these systems is based on cascading grammars and the theory of finite-state automata. One of the earliest languages for specifying such rules is the Common Pattern Specification Language (CPSL) developed in the context of the TIPSTER project [2]. To overcome some of the drawbacks of CPSL resulting from a sequential view of the input, the AfST system [4] uses a more powerful grammar that views its input as an object graph. Beginning in the mid-1990s, as the unstruc-tured content on the Web continued to grow, information extraction techniques were applied in building popular Web applications. One of the earliest such uses of information extraction was in the context of screen scraping for on-line comparison shopping and data integration applications. By manually examining a number of sample pages, application designers would de-velop ad hoc rules and regular expressions to eke out relevant pieces of information (e.g., the name
Cite
CITATION STYLE
Chiticariu, L., Danilevsky, M., Ho, H., Krishnamurthy, R., Li, Y., Raghavan, S., … Zhu, H. (2016). Web Information Extraction. In Encyclopedia of Database Systems (pp. 1–9). Springer New York. https://doi.org/10.1007/978-1-4899-7993-3_459-2
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.