Abstract
Variable selection is important for developing accurate and interpretable prediction models. While classical and penalized methods are widely used, few simulation studies provide meaningful comparisons. This study compares their predictive performance and model complexity in low-dimensional data. Three classical methods (best subset selection, backward elimination, and forward selection) and four penalized methods (nonnegative garrote (NNG), lasso, adaptive lasso (ALASSO), and relaxed lasso (RLASSO)) were compared. Tuning parameters were selected using cross-validation (CV), Akaike information criterion (AIC), and Bayesian information criterion (BIC). Classical methods performed similarly and produced worse predictions than penalized methods in limited-information scenarios (small samples, high correlation, and low signal-to-noise ratio (SNR)), but performed comparably or better in sufficient-information scenarios (large samples, low correlation, and high SNR). Lasso was superior under limited information but was less effective in sufficient-information scenarios. NNG, ALASSO, and RLASSO outperformed lasso in sufficient-information scenarios, with no clear winner among them. AIC and CV produced similar results and outperformed BIC, except in sufficient-information settings, where BIC performed better. Our findings suggest that no single method consistently outperforms others, as performance depends on the amount of information in the data. Lasso is preferred in limited-information settings, whereas classical methods are more suitable in sufficient-information settings, as they also tend to select simpler models.
Author supplied keywords
Cite
CITATION STYLE
Kipruto, E., & Sauerbrei, W. (2025). Evaluating Prediction Performance: A Simulation Study Comparing Penalized and Classical Variable Selection Methods in Low-Dimensional Data. Applied Sciences (Switzerland), 15(13). https://doi.org/10.3390/app15137443
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.