Learning Document Similarity Using Natural Language Processing

  • Merlo P
  • Henderson J
  • Schneider G
  • et al.
N/ACitations
Citations of this article
32Readers
Mendeley users who have this article in their library.

Abstract

The recent considerable growth in the amount of easily available on-line text has brought to the foreground the need for large-scale natural language processing tools for text data mining. In this paper we address the problem of organizing documents into meaningful groups according to their content and to visualize a text collection, providing an overview of the range of documents and of their relationships, so that they can be browsed more easily. We use Self-Organizing Maps (SOMs) (Kohonen 1984). Great efficiency challenges arise in creating these maps. We study linguistically-motivated ways of reducing the representation of a document to increase efficiency and ways to disambiguate the words in the documents.

Cite

CITATION STYLE

APA

Merlo, P., Henderson, J., Schneider, G., & Wehrli, E. (2003). Learning Document Similarity Using Natural Language Processing. Linguistik Online, 17(5). https://doi.org/10.13092/lo.17.788

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free