Classification performance bias between training and test sets in a limited mammography dataset

Rui Hou; Joseph Y. Lo; Jeffrey R. Marks; E. Shelley Hwang; Lars J. Grimm

Journal ArticleOPEN ACCESS

Classification performance bias between training and test sets in a limited mammography dataset

PLoS ONE (2024) 19(2 February)

DOI: 10.1371/journal.pone.0282402

2Citations

18Readers

Get full text

Abstract

Objectives To assess the performance bias caused by sampling data into training and test sets in a mammography radiomics study. Methods Mammograms from 700 women were used to study upstaging of ductal carcinoma in situ. The dataset was repeatedly shuffled and split into training (n = 400) and test cases (n = 300) forty times. For each split, cross-validation was used for training, followed by an assessment of the test set. Logistic regression with regularization and support vector machine were used as the machine learning classifiers. For each split and classifier type, multiple models were created based on radiomics and/or clinical features. Results Area under the curve (AUC) performances varied considerably across the different data splits (e.g., radiomics regression model: train 0.58–0.70, test 0.59–0.73). Performances for regression models showed a tradeoff where better training led to worse testing and vice versa. Cross-validation over all cases reduced this variability, but required samples of 500+ cases to yield representative estimates of performance. Conclusions In medical imaging, clinical datasets are often limited to relatively small size. Models built from different training sets may not be representative of the whole dataset. Depending on the selected data split and model, performance bias could lead to inappropriate conclusions that might influence the clinical significance of the findings.

Cite

CITATION STYLE

APA

Hou, R., Lo, J. Y., Marks, J. R., Hwang, E. S., & Grimm, L. J. (2024). Classification performance bias between training and test sets in a limited mammography dataset. PLoS ONE, 19(2 February). https://doi.org/10.1371/journal.pone.0282402

Classification performance bias between training and test sets in a limited mammography dataset

Abstract

Cite

Register to see more suggestions