Data quality can be seen as a very important factor for the validity of information extracted from data sets using statistical or data mining procedures. In the paper we propose a description of data quality allowing us to characterize data quality of the whole data set, as well as data quality of particular variables and individual cases. On the basis of the proposed description, we define a distance based measure of data quality for individual cases as a distance of the cases from the ideal one. Such a measure can be used as additional information for preparation of a training data set, fitting models, decision making based on results of analyses etc. It can be utilized in different ways ranging from a simple weighting function to belief functions.
CITATION STYLE
Král, P., Sobíšek, L., & Stachová, M. (2014). A distance based measure of data quality. Metodoloski Zvezki, 11(2), 107–120. https://doi.org/10.51936/npie9973
Mendeley helps you to discover research relevant for your work.