3 problems
- 0 votes0 replies0 views
Worse dependence or non-convergence of AdaGrad under infinite-variance noise
Consider AdaGrad applied to smooth convex optimization problems with stochastic gradient noise whose bounded -th moment satisfies , and let denote the adapt…
- 0 votes0 replies1 view
High-probability convergence of Adam under heavy-tailed gradient noise
Stochastic optimization methods such as Adam and Clip-SGD use adaptive stepsizes, and Clip-SGD is known to converge in expectation when the gradient noise has a bounded -th…
- 0 votes0 replies0 views
Improved logarithmic factors for high-probability stochastic minimization
Logarithmic-factor improvement conjecture. Adjusting the proof technique from Nguyen et al. should improve the logarithmic factors in the authors' results as well.