Abstract
The paper aims to extend the theory and application of nonconvex Newton-type methods, namely trust region and cubic regularization, to the settings in which, in addition to the solution of subproblems, the gradient and the Hessian of the objective function are approximated. Using certain conditions on such approximations, the paper establishes optimal worst-case iteration complexities as the exact counterparts. This paper is part of a broader research program on designing, analyzing, and implementing efficient second-order optimization methods for large-scale machine learning applications. The authors were based at UC Berkeley when the idea of the project was conceived. The first two authors were PhD students, the third author was a postdoc, all supervised by the fourth author.For solving large-scale nonconvex problems, we propose inexact variants of trust region and adaptive cubic regularization methods, which, to increase efficiency, incorporate various approximations. In particular, in addition to inexact subproblem solves, both the gradient and Hessian are suitably estimated. Using certain conditions on such approximations, we show that our proposed inexact methods achieve similar optimal worst-case iteration complexities as the exact counterparts. In the context of finite-sum problems, we then explore randomized subsampling methods as ways to construct the gradient and Hessian approximations and examine the empirical performance of our algorithms on some model problems. We empirically demonstrate that our proposed algorithms are practically implementable in that failure to precisely fine-tune the associated hyperparameters is unlikely to result in unwanted behaviors, for example, divergence or stagnation.
Cite
CITATION STYLE
Yao, Z., Xu, P., Roosta, F., & Mahoney, M. W. (2021). Inexact Nonconvex Newton-Type Methods. INFORMS Journal on Optimization, 3(2), 154–182. https://doi.org/10.1287/ijoo.2019.0043
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.