A distance based measure of data quality

Pavol Král; Lukáš Sobíšek; Mária Stachová

Journal ArticleOPEN ACCESS

A distance based measure of data quality

Metodoloski Zvezki (2014) 11(2) 107-120

DOI: 10.51936/npie9973

0Citations

6Readers

Abstract

Data quality can be seen as a very important factor for the validity of information extracted from data sets using statistical or data mining procedures. In the paper we propose a description of data quality allowing us to characterize data quality of the whole data set, as well as data quality of particular variables and individual cases. On the basis of the proposed description, we define a distance based measure of data quality for individual cases as a distance of the cases from the ideal one. Such a measure can be used as additional information for preparation of a training data set, fitting models, decision making based on results of analyses etc. It can be utilized in different ways ranging from a simple weighting function to belief functions.

Cite

CITATION STYLE

APA

Král, P., Sobíšek, L., & Stachová, M. (2014). A distance based measure of data quality. Metodoloski Zvezki, 11(2), 107–120. https://doi.org/10.51936/npie9973

A distance based measure of data quality

Abstract

Cite

Register to see more suggestions