Abstract
While articles assessing the accuracy of traditional statistical packages are fairly commonplace, data mining software has escaped this important scrutiny. We apply the National Institute of Standards and Technology Statistical Reference Datasets tests for the numerical accuracy of statistical packages to 7 data mining packages: IBM Modeler, KNIME, Orange, Python, RapidMiner, Weka, and XLMiner. We find that one package has an unstable algorithm for the calculation of the sample variance and only two have reliable linear regression routines. Of these two packages that offer analysis of variance, one has a bad algorithm. The accuracy of statistical calculations in data mining packages cannot be taken for granted. This article is categorized under: Technologies > Statistical Fundamentals Algorithmic Development > Statistics Application Areas > Data Mining Software Tools.
Author supplied keywords
Cite
CITATION STYLE
McCullough, B. D., Mokfi, T., & Almaeenejad, M. (2019). On the accuracy of linear regression routines in some data mining packages. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9(3). https://doi.org/10.1002/widm.1279
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.