Customized Information Extraction as a Basis for Resource Discovery

9Citations
Citations of this article
19Readers
Mendeley users who have this article in their library.

Abstract

Indexing file contents is a powerful means of helping users locate documents, software, and other types of data among large repositories. In environments that contain many different types of data, content indexing requires type-specific processing to extract information effectively. We present a model for type-specific, user-customizable information extraction, and a system implementation called Essence. This software structure allows users to associate specialized extraction methods with ordinary files, providing the illusion of an object-oriented file system that encapsulates indexing methods within files. By exploiting the semantics of common file types, Essence generates compact yet representative file summaries that can be used to improve both browsing and indexing in resource discovery systems. Essence can extract information from most of the types of files found in common file systems, including files with nested structure (such as compressed "tar" files). Essence interoperates with a number of different search/index systems (such as WAIS and Glimpse), as part of the Harvest system.

Cite

CITATION STYLE

APA

Hardy, D. R., & Schwartz, M. F. (1996). Customized Information Extraction as a Basis for Resource Discovery. ACM Transactions on Computer Systems, 14(2), 171–199. https://doi.org/10.1145/227695.227697

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free