Automatically generating data linkages using a domain-independent candidate selection approach

Dezhao Song; Jeff Heflin

Conference ProceedingsOPEN ACCESS

Automatically generating data linkages using a domain-independent candidate selection approach

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2011) 7031 LNCS(PART 1) 649-664

DOI: 10.1007/978-3-642-25073-6_41

71Citations

35Readers

Get full text

Abstract

One challenge for Linked Data is scalably establishing high-quality owl:sameAs links between instances (e.g., people, geographical locations, publications, etc.) in different data sources. Traditional approaches to this entity coreference problem do not scale because they exhaustively compare every pair of instances. In this paper, we propose a candidate selection algorithm for pruning the search space for entity coreference. We select candidate instance pairs by computing a character-level similarity on discriminating literal values that are chosen using domain-independent unsupervised learning. We index the instances on the chosen predicates' literal values to efficiently look up similar instances. We evaluate our approach on two RDF and three structured datasets. We show that the traditional metrics don't always accurately reflect the relative benefits of candidate selection, and propose additional metrics. We show that our algorithm frequently outperforms alternatives and is able to process 1 million instances in under one hour on a single Sun Workstation. Furthermore, on the RDF datasets, we show that the entire entity coreference process scales well by applying our technique. Surprisingly, this high recall, low precision filtering mechanism frequently leads to higher F-scores in the overall system. © 2011 Springer-Verlag.

Author supplied keywords

Cite

CITATION STYLE

APA

Song, D., & Heflin, J. (2011). Automatically generating data linkages using a domain-independent candidate selection approach. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 7031 LNCS, pp. 649–664). https://doi.org/10.1007/978-3-642-25073-6_41

Automatically generating data linkages using a domain-independent candidate selection approach

Abstract

Author supplied keywords

Cite

Register to see more suggestions