Data-Centric Machine Learning: Improving Model Performance and Understanding through Dataset Analysis

11Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.

Abstract

Machine learning research typically starts with a fixed data set created early in the process. The focus of the experiments is finding a model and training procedure that result in the best possible performance in terms of some selected evaluation metric. This paper explores how changes in a data set influence the measured performance of a model. Using three publicly available data sets from the legal domain, we investigate how changes to their size, the train/test splits, and the human labelling accuracy impact the performance of a trained deep learning classifier. Our experiments suggest that analyzing how data set properties affect performance can be an important step in improving the results of trained classifiers, and leads to better understanding of the obtained results.

Cite

CITATION STYLE

APA

Westermann, H., Šavelka, J., Walker, V. R., Ashley, K. D., & Benyekhlef, K. (2021). Data-Centric Machine Learning: Improving Model Performance and Understanding through Dataset Analysis. In Frontiers in Artificial Intelligence and Applications (Vol. 346, pp. 54–57). IOS Press BV. https://doi.org/10.3233/FAIA210316

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free