79 problems
- 0 votes0 replies0 views
Saddle-point avoidance conjecture for backtracking gradient descent
Let be a function that is near its critical points, and let be the sequence generated by the Backtracking GD method from…
- 0 votes0 replies1 view
Optimality of the silver stepsize rate among non-adaptive schedules
Silver-rate optimality conjecture. The asymptotic rate achieved by the silver stepsize schedule is optimal among all non-adaptive stepsize schedules.
- 0 votes0 replies0 views
Asymptotic optimality of the silver convergence rate for stepsize schedules
For minimizing convex quadratics with gradient descent using stepsize schedules, the silver convergence rate is . Silver-rate optimality conjecture. T…
- 0 votes0 replies0 views
Fixed-step-size convergence conjecture for Acc-DNGD-NSC
Fixed-step-size convergence conjecture. The algorithm should converge with rate
- 0 votes0 replies0 views
Parabolic strong stable foliation conjecture near a manifold of flat minima
Let be the invariant set and let be the local invariant manifold appearing in the construction, with coordinates . The -axes are the d…
- 0 votes0 replies0 views
Polylogarithmic data-dimension dependence for LPI-GD
Let denote the data dimension and the number of training samples in local polynomial interpolation-based gradient descent (LPI-GD). Existing guarantees assume…
- 0 votes0 replies0 views
Exact optimality of the silver stepsize schedule in strongly convex optimization
Silver-schedule optimality conjecture. The silver stepsize schedule is exactly optimal for this performance metric among stepsize schedules in the strongly convex setting.
- 0 votes0 replies0 views
Global convergence for wide shallow models with vector output parameters
Global convergence conjecture. The property holds when .
- 0 votes0 replies1 view
Non-singularity conjecture for gradient descent maps of recurrent neural networks
Let a recurrent neural network have loss landscape and associated gradient descent (GD) map. Non-singularity conjecture. The GD map is non-singular for the loss landscape of a recu…
- 0 votes0 replies0 views
Rapid convergence to a tubular neighbourhood during progressive sharpening
Let GD denote gradient descent, and let the solution manifold be the set of global minimisers of the overparametrised least-squares problem. A tubular neighbourhood is a neighbourh…
- 0 votes0 replies0 views
Sarao's gradient-flow convergence conjecture for Wirtinger flow
Let be the planted parameter, let be the loss associated with the Wirtinger-flow objective … and consider the proportional regime .…
- 0 votes0 replies1 view
Sparsity conjecture for the support of the optimal coupling
In the setting of quadratically regularized optimal transport, let the support of the optimal coupling be partitioned into sections, and let denote the underlying dimension and…
- 0 votes0 replies0 views
Positive growth-rate conditions for robust gradient descent beyond
Positive growth-rate conjecture. Positive growth-rate conditions on the stepsizes that allow should exist.
- 0 votes0 replies0 views
Optimized full inexact gradient descent accelerates functional accuracy under inexactness
Let denote the iteration count, let be the smoothness constant, and consider inexact gradient descent for minimizing under . Let…
- 0 votes0 replies1 view
Optimized full inexact gradient descent has a superlinear numerical rate for residual gradient minimization
Let denote the iteration count and let the performance criterion be the minimum squared gradient norm, , for minimizing a smooth c…
- 0 votes0 replies1 view
The unit-sphere global convergence conjecture for logistic regression
Let the data lie on the unit sphere, and let be the stability-threshold step size, where is the step-size parameter and is the relevant s…
- 0 votes0 replies1 view
Beyond-energy-dissipation conjecture for Allen–Cahn time-step restrictions
Consider the gradient descent scheme with convex–concave splitting for the Allen–Cahn equation, and its energy-dissipation estimate, which includes the energy decrease and a neglec…
- 0 votes0 replies0 views
Geometric lower-semicontinuity conjecture for Allen–Cahn interface movement
Geometric lower-semicontinuity conjecture. A geometric condition on could be used to improve the Hölder exponent in the estimate for the displacement of the iterate…
- 0 votes0 replies0 views
Tightness of the intermediate regime for relatively inexact gradient descent
Intermediate-regime tightness conjecture. The intermediate regime in Theorem 1 is tight. Furthermore, the corresponding worst-case function is bivariate.
- 0 votes0 replies0 views
Noise robustness of benign overfitting in the weak signal regime
Consider the weak signal regime for gradient descent in leaky ReLU two-layer neural networks, and introduce noise to the binary-classification labels. Noise-robustness conjecture.…
- 0 votes0 replies1 view
Minimal activation regularity conjecture for consistent test-error estimation
Let be the activation functions used in the neural network, and let and denote their first and second derivatives. Theorem establishe…
- 0 votes0 replies0 views
Minimal regularity conjecture for gradient descent state evolution
The neural network has layers indexed by , with activation functions . Theorem establishes the stated gradient descent dynamics under the paper's smoot…
- 0 votes0 replies0 views
Grimmer's ratio-property conjecture for PEP multipliers
Grimmer's ratio-property conjecture. There is a set of multipliers satisfying this ratio property whenever , and this set proves T…
- 0 votes0 replies0 views
Grimmer's structured-multiplier conjecture for the exact gradient descent rate
Let be the number of gradient descent iterations, let and denote the function parameters, let be the optimal stepsize, and let be the…
- 0 votes0 replies0 views
Basic g-composable schedules capture minimax-optimal gradient-norm schedules
Basic g-composable schedule conjecture. For each , every minimax-optimal stepsize schedule solving