Abstract
We study a localized notion of uniform convergence known as an “optimistic rate” [ 34 , 39 ] for linear regression with Gaussian data. Our refined analysis avoids the hidden constant and logarithmic factor in existing results, which are known to be crucial in high-dimensional settings, especially for understanding interpolation learning. As a special case, our analysis recovers the guarantee from Koehler et al. [ 21 ], which tightly characterizes the population risk of low-norm interpolators under the benign overfitting conditions. Our optimistic rate bound, though, also analyzes predictors with arbitrary training error. This allows us to recover some classical statistical guarantees for ridge and LASSO regression under random designs and helps us obtain a precise understanding of the excess risk of near-interpolators in the over-parameterized regime. Problem Statement Generalization theory proposes to explain the ability of machine learning models to generalize to fresh examples by bounding the gap between the test error (error on new examples) and training error (error on the data they were trained upon). Most generalization bounds are too loose to explain the performance of overfit models which nevertheless generalize well. Recently proposed bounds using “uniform convergence of interpolators” can explain this phenomena in the context of linear models, but such bounds are specific to heavily overfit models. Are there generalization bounds which can provide a unified picture for both overfit and heavily regularized models? Methods We study this problem in the setting of linear models with Gaussian covariates and the squared loss. At a technical level, this allows us to analyze the (nonconvex) generalization landscape using tools from Gaussian processes, in particular Gordon's theorem, which has proven very powerful in the analysis of M-estimation in previous works. Results We establish a sharp generalization bound which fits cleanly into an existing framework in the statistical learning theory literature: optimistic rates theory. Our bound is nonasymptotic and controls the generalization gap by a function of the training error and Rademacher complexity of a class of functions. Unlike previous optimistic rates bounds, our bound has sharp constants and we show it can explain the ability of heavily overfit predictors to generalize well (“benign overfitting”), in particular recovering results “uniform convergence of interpolators” as a special case. We show that our result can cleanly recover many other results concerning the performance of M-estimators like the LASSO and ridge regression. Significance This result gives insights into what optimal generalization bounds look like and how they reconcile modern phenomena like overfitting and double descent with the mathematical theory of statistical learning. It also establishes connections between tools like Rademacher complexity from learning theory and exact proportional asymptotics studied in high-dimensional statistics and other areas.
Cite
CITATION STYLE
Zhou, L., Koehler, F., Sutherland, D. J., & Srebro, N. (2024). Optimistic Rates: A Unifying Theory for Interpolation Learning and Regularization in Linear Regression. ACM / IMS Journal of Data Science, 1(2), 1–51. https://doi.org/10.1145/3594234
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.