Text categorization is usually performed by supervised algorithms on the large amount of hand-labelled documents which are labor-intensive and often not available. To avoid this drawback, this paper proposes a text categorization approach that is designed to fully exploiting semantic resources. It employs the ontological knowledge not only as lexical support for disambiguating terms and deriving their sense inventory, but also to classify documents in topic categories. Specifically, our work relates to apply two corpus-based thesauri (i.e. WordNet and WordNet Domains) for selecting the correct sense of words in a document while utilizing domain names for classification purposes. Experiments presented show how our approach performs well in classifying a large corpus of documents. A key part of the paper is the discussion of important aspects related to the use of surrounding words and different methods for word sense disambiguation. © 2013 Springer International Publishing.
CITATION STYLE
Dessì, N., Dessì, S., & Pes, B. (2014). A fully semantic approach to large scale text categorization. In Lecture Notes in Electrical Engineering (Vol. 264 LNEE, pp. 149–157). Springer Verlag. https://doi.org/10.1007/978-3-319-01604-7_15
Mendeley helps you to discover research relevant for your work.