Incremental Unsupervised Name Disambiguation in Cleaned Digital Libraries

  • de Carvalho A
  • Ferreira A
  • Laender A
  • et al.
ISSN: 21787107
N/ACitations
Citations of this article
33Readers
Mendeley users who have this article in their library.

Abstract

Name ambiguity in the context of bibliographic citations is one of the hardest problems currently faced by the Digital Library (DL) community. Here we deal with the problem of disambiguating new citations records inserted into a cleaned DL, without the need to process the whole collection, which is usually necessary for unsupervised methods. Although supervised solutions can deal with this situation, there is the costly burden of generating training data besides the fact that these methods cannot handle well the insertion of records of new authors not already existent in the repository. In this article, we propose a new unsupervised method that identifies the correct authors of the new citation records to be inserted in a DL. The method is based on heuristics that are also used to identify whether the new records belong to authors already in the digital library or not, correctly identifying new authors in most cases. Our experimental evaluation, using synthetic and real datasets, shows gains of up to 19\% when compared to a state-of-the-art method without the cost of having to disambiguate the whole DL at each new load (as done by unsupervised methods) or the need for any training (as done by supervised methods).

Cite

CITATION STYLE

APA

de Carvalho, A. P., Ferreira, A. A., Laender, A. H. F., & Gonçalves, M. A. (2011). Incremental Unsupervised Name Disambiguation in Cleaned Digital Libraries. Journal of Information and Data Management, 2(573871), 289.

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free