Abstract
Vast amounts of text on the Web are unstructured and ungrammatical, such as clas- sified ads, auction listings, forum postings, etc. We call such text \posts." Despite their inconsistent structure and lack of grammar, posts are full of useful information. This pa- per presents work on semi-automatically building tables of relational information, called \reference sets," by analyzing such posts directly. Reference sets can be applied to a num- ber of tasks such as ontology maintenance and information extraction. Our reference-set construction method starts with just a small amount of background knowledge, and con- structs tuples representing the entities in the posts to form a reference set. We also describe an extension to this approach for the special case where even this small amount of back- ground knowledge is impossible to discover and use. To evaluate the utility of the machine- constructed reference sets, we compare them to manually constructed reference sets in the context of reference-set-based information extraction. Our results show the reference sets constructed by our method outperform manually constructed reference sets. We also com- pare the reference-set-based extraction approach using the machine-constructed reference set to supervised extraction approaches using generic features. These results demonstrate that using machine-constructed reference sets outperforms the supervised methods, even though the supervised methods require training data. © 2010 AI Access Foundation.
Cite
CITATION STYLE
Michelson, M., & Knoblock, C. A. (2010). Constructing reference sets from unstructured, ungrammatical text. Journal of Artificial Intelligence Research, 38, 189–221. https://doi.org/10.1613/jair.2937
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.