7 problems
- 0 votes0 replies1 view
The learning-rate-annealing conjecture for ScheduleFree diffusion training
Learning-rate-annealing conjecture. The mismatch between similar loss values and generative quality under ScheduleFree is partially due to missing learning-rate annealing: adding a…
- 0 votes0 replies0 views
Conjecture on asymptotic quadratic behavior behind the success of large learning rates
In the quadratic case, large fixed step-sizes achieve optimal convergence rates. Asymptotic quadratic-behavior conjecture. The success of large learning rates may be attributed to…
- 0 votes0 replies0 views
Conjecture that calibrated learning rates retain desirable concentration rates
A Gibbs posterior uses a learning rate to control the contribution of the loss relative to the prior; in the calibration approach of Syring and Martin, the learning rate is obtaine…
- 0 votes0 replies0 views
Learning-rate trade-off conjecture for local update methods
Learning-rate trade-off conjecture. Even in much broader settings, the choice of learning rate dictates a trade-off between accuracy and initial convergence.
- 0 votes0 replies0 views
Learning-rate relaxation conjecture for local update methods
Learning-rate relaxation conjecture. This condition can be relaxed. The question concerns whether convergence guarantees under a bounded variance assumption can hold with a larger…
- 0 votes0 replies0 views
Final-iterate convergence conjecture for exponentially decaying learning rates
The algorithm of Allen-Zhu works with iterate averaging, rather than the final iterate, and uses an exponentially decaying learning-rate scheme. Final-iterate convergence conjectur…
- 0 votes0 replies0 views
Adaptive learning-rate convergence conjecture for smoothed stochastic gradient descent
Adaptive learning-rate conjecture. Choosing the learning rate adaptively according to the size of the current gradient might improve the convergence requirement to