Validation of Deduplication in Data using Similarity Measure

  • Wandhekar V
  • Mohanpurkar A
N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Deduplication is the process of determining all categories of information within a data set that signify the same real life / world entity. The data gathered from various resources may have data high quality issues in it. The concept to identify duplicates by using windowing and blocking strategy. The objective is to achieve better precision, good efficiency and also to reduce the false positive rate all are in accordance with the estimated similarities of records. Various Similarity metrics are commonly used to recognize the similar field entries. So the main focus of this paper is to applying appropriate similarity measure on appropriate data to properly identifying the duplicates.

Cite

CITATION STYLE

APA

Wandhekar, V., & Mohanpurkar, A. (2015). Validation of Deduplication in Data using Similarity Measure. International Journal of Computer Applications, 116(21), 18–22. https://doi.org/10.5120/20460-2819

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free