54 problems
- 0 votes0 replies0 views
Improved convergence rate under higher moments for stochastic-gradient noise
Higher-moment rate conjecture. The rate can be improved by assuming that has a high-order moment.
- 0 votes0 replies0 views
Yun–Sra–Jadbabaie SS–RS–GD inequalities
Yun–Sra–Jadbabaie conjecture. There exists a constant such that, whenever for every , one has
- 0 votes0 replies0 views
Jain et al.'s anytime optimality conjecture for the last iterate of SGD
Jain et al.'s anytime optimality conjecture. In the absence of a priori information about , no stepsize sequence can ensure the information-theoretically optimal error rate for…
- 0 votes0 replies0 views
Conjecture on bounded gains from lookahead optimization
Bounded-gains conjecture. Lookahead can provide modest but bounded gains in more general settings.
- 0 votes0 replies0 views
Steady-state convergence and moment conjecture for general-objective SGD
Let be convex and of class for an even integer , with unique minimizer , satisfying for…
- 0 votes0 replies1 view
Maximal-power conjecture for the transition in multi-class stochastic gradient descent
The setting is the two-class stochastic-gradient-descent model discussed above, with power-law covariance and mean parameters whose exponents include ; the relevant transit…
- 0 votes0 replies1 view
Logarithm-free last-iterate convergence for convex smooth stochastic optimization
Logarithm-free last-iterate conjecture. It should be possible to eliminate the term from the best possible last-iterate convergence bound for convex smooth problems; in par…
- 0 votes0 replies1 view
Conjecture that the sieve SGD proof requirement can be improved
The discussion concerns sieve stochastic gradient descent estimators with tuning parameters , , and , where the current stability result requires a condition of…
- 0 votes0 replies0 views
The Lyapunov-framework conjecture on the optimal SGD bias bound
Lyapunov-framework conjecture. Within this Lyapunov framework, it is not possible to obtain a bound for SGD with a bias term of order
- 0 votes0 replies0 views
Extension of the summarized scaling law to time-dependent momentum parameters
The summarized scaling law theorem concerns the scaling behavior established earlier in the paper for the considered momentum algorithms and their dimension-dependent hyperparamete…
- 0 votes0 replies0 views
Conjecture on the slowest-eigenvalue dominance of bias-error decay
Consider stochastic gradient descent with exponential moving average, and let the bias error be decomposed into components associated with the eigenspaces of the data-feature covar…
- 0 votes0 replies0 views
Global localized approximation conjecture for stochastic gradient descent
Let have local minima , and let be the solution of the localized equation around the local minimum , for . Let denote the s…
- 0 votes0 replies0 views
Conjecture on the interrelations and plateau timings of SGD sharpness metrics
Let stochastic gradient descent (SGD) denote the mini-batch optimization dynamics, and consider the metrics , Batch Sharpness, and Interaction-Aware Sharpness (IAS)…
- 0 votes0 replies0 views
Conjecture on batch sharpness reaching the GNI threshold through gradient-noise variation
The conjecture. Batch Sharpness reaches at by means of increasing more than the alignment of gradients wi…
- 0 votes0 replies0 views
Implicit-bias conjecture for stochastic gradient descent in equivariant networks
The hypothesis space consists of functions represented by the single-hidden-layer equivariant networks studied in the paper, and norm-bounded functions form a specified subset of t…
- 0 votes0 replies0 views
Instance-dependent weighting can improve schedule-free optimization
The framework uses exponentially increasing weights, corresponding to a constant , and these weights achieve optimal worst-case convergence guarantees. Instance-depende…
- 0 votes0 replies0 views
Abbe et al.'s leap-complexity conjecture for online SGD
Let be a Boolean function, and let its leap complexity be the minimum value for which there is an ordering of its coeffic…
- 0 votes0 replies0 views
The conjecture on computational costs when
Computational-cost conjecture. More computational costs are essentially required under the regime .
- 0 votes0 replies0 views
Strong-regularization explanation for the clean neural scaling law
In an infinite-dimensional linear regression model, suppose only -dimensional sketched covariates are observed, a linear predictor with trainable parameters is trained by on…
- 0 votes0 replies0 views
Conjecture on the optimal lambda-dependence of SME-2 and SPF error bounds
The stochastic-gradient methods and their SDE approximations are considered under the smoothness and convexity assumptions stated for the weak approximation theorem above. In parti…
- 0 votes0 replies0 views
Conjecture on the deterministic equivalent for SGD risk
Deterministic-equivalent conjecture. For all admissible and , with probability tending to as ,
- 0 votes0 replies1 view
A conjectural extension of centered-noise analyses to the de-biased trajectory
Let mini-batch SGD without replacement be compared with an algorithm whose steps are centered and independent, such as full-batch gradient descent or SGD with replacement. The depe…
- 0 votes0 replies0 views
A diffusion-powered escape conjecture for SGD without replacement
The setting concerns stochastic gradient descent (SGD) without replacement, in which the data batches used at successive steps are dependent because they are disjoint, and saddle e…
- 0 votes0 replies0 views
Eventual one-sided margin conjecture for oscillating stochastic gradient descent
Let be the model output at iteration , let be the label, and let and \underaccent{\bar}{T}_k denote the corresponding stopping…
- 0 votes0 replies0 views
Path differentiability conjecture for the sliced Wasserstein loss
Let , and define the sliced Wasserstein energy by … Here is the one-dimensional projected Wasserstein loss between and , and path dif…