A distance based measure of data quality

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Data quality can be seen as a very important factor for the validity of information extracted from data sets using statistical or data mining procedures. In the paper we propose a description of data quality allowing us to characterize data quality of the whole data set, as well as data quality of particular variables and individual cases. On the basis of the proposed description, we define a distance based measure of data quality for individual cases as a distance of the cases from the ideal one. Such a measure can be used as additional information for preparation of a training data set, fitting models, decision making based on results of analyses etc. It can be utilized in different ways ranging from a simple weighting function to belief functions.

Cite

CITATION STYLE

APA

Král, P., Sobíšek, L., & Stachová, M. (2014). A distance based measure of data quality. Metodoloski Zvezki, 11(2), 107–120. https://doi.org/10.51936/npie9973

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free