14 problems
- 0 votes0 replies1 view
Gunasekar et al.'s minimum nuclear norm conjecture for full-dimensional matrix factorization
Let be a full-dimensional Burer--Monteiro factorization, with gradient descent applied using sufficiently small step sizes and initialization sufficiently close to the…
- 0 votes0 replies0 views
Implicit-bias conjecture for stochastic gradient descent in equivariant networks
The hypothesis space consists of functions represented by the single-hidden-layer equivariant networks studied in the paper, and norm-bounded functions form a specified subset of t…
- 0 votes0 replies0 views
Adam's -geometry conjecture for outperforming gradient descent
Adam's -geometry conjecture. outperforms due to its utilization of geometry, under which the loss function could have bette…
- 0 votes0 replies0 views
Conjecture that deep models cannot be collapsed for depth greater than two
Consider deep models with depth parameter . Collapse conjecture. We conjecture that such deep models cannot be "collapsed". The source gives only a heuristic argument and no f…
- 0 votes0 replies0 views
Rigidity from uniqueness of the minimum norm interpolant
Consider the radially symmetric setting studied for the network interpolation problem, and let a minimum norm interpolant be a function that fits the prescribed data while minimizi…
- 0 votes0 replies0 views
Rigidity from uniqueness of the radially symmetric minimum norm interpolant
The setting is a radially symmetric interpolation problem for a ReLU network, with minimum norm interpolants sought among functions fitting the prescribed data. In the one-dimensio…
- 0 votes0 replies0 views
The bad-regularity conjecture for neural-network learning rates
Consider neural-network training objectives with good or bad regularity, trained using sufficiently large learning rates. The relevant large-learning-rate phenomena include edge of…
- 0 votes0 replies0 views
The extension conjecture to minima at
Consider the function class analyzed in the paper, whose global minimizers are fixed at , with the parameter restriction that excludes even values of the relevant exponent. E…
- 0 votes0 replies0 views
The global-regularity conjecture for large learning-rate phenomena
Let a function be studied through large-learning-rate phenomena such as edge of stability and balancing, and distinguish global regularity properties from regularity observed only…
- 0 votes0 replies0 views
The regularity conjecture for large learning-rate phenomena
Consider objective functions trained by gradient descent with large learning rates, and measure their regularity by an appropriate regularity notion. The phenomena of interest incl…
- 0 votes0 replies0 views
The flat-minima conjecture for large learning rates
Large learning rates are considered in the context of gradient-based training, where a flat minimum is a minimum associated with favorable generalization properties. Flat-minima co…
- 0 votes0 replies0 views
The Adam momentum–second-moment conjecture for test error
Let and denote the momentum and second-moment decay parameters, respectively, in Adam. Consider a stable regime of training, and let the test error be the resulting…
- 0 votes0 replies0 views
Gunasekar's implicit-bias conjecture for non-commuting matrix factorizations
Let be a matrix-factorization parametrization, such as or , and distinguish diagonal measurements, whose coordinate parametrizations commute, fro…
- 0 votes0 replies0 views
Gradient Starvation's protection against overfitting
In the setting where training and test data points are drawn from the same distribution, the most salient features are predictive in both sets. Gradient Starvation's overfitting co…