31 problems
- 0 votes0 replies0 views
The gradient-accumulation generalization conjecture
Gradient accumulation uses large-batch samples for gradient evaluation. Gradient-accumulation generalization conjecture. Although gradient accumulation can help during optimization…
- 0 votes0 replies0 views
Unique-extension generalization conjecture for k × n Chomp
For each positive integer , a P-position of Chomp consists of a weakly decreasing sequence of row lengths. Unique-extension generalization conjecture. If the Unique…
- 0 votes0 replies1 view
Broader regression-function conjecture for consistent test-error estimation
Consider the estimator constructed by Algorithm; Theorem proves its consistency under the paper's single-index or multi-index reg…
- 0 votes0 replies1 view
Minimal activation regularity conjecture for consistent test-error estimation
Let be the activation functions used in the neural network, and let and denote their first and second derivatives. Theorem establishe…
- 0 votes0 replies0 views
NASM's architectural explanation for ID/OOD generalization
NASM is a neural adaptive spectral method whose architecture explicitly disentangles coefficients from basis functions. NASM generalization conjecture. We conjecture that NASM's ab…
- 0 votes0 replies0 views
The RASP length-generalization conjecture for transformers
RASP length-generalization conjecture. Based on extensive empirical results, transformers tend to length-generalize on tasks that can be solved by RASP.
- 0 votes0 replies0 views
Extension of the weak generalization hypothesis to general architectures
A modular polynomial on in two variables is said to satisfy the weak generalization hypothesis when it has the form … where , , and are…
- 0 votes0 replies0 views
The conjecture that training-data reuse for recalibration leads to severe overfitting
Let a model be recalibrated using the same training data on which the model was trained, and let the resulting calibration be evaluated by an expected calibration error or related…
- 0 votes0 replies0 views
The flat-minima conjecture for large learning rates
Large learning rates are considered in the context of gradient-based training, where a flat minimum is a minimum associated with favorable generalization properties. Flat-minima co…
- 0 votes0 replies0 views
Loshchilov–Hutter conjecture on quadratic regularization and neural-network generalization
Consider training neural networks with quadratic regularization terms. Loshchilov–Hutter conjecture. The quadratic regularization terms contribute to the low generalization error i…
- 0 votes0 replies0 views
The Adam momentum–second-moment conjecture for test error
Let and denote the momentum and second-moment decay parameters, respectively, in Adam. Consider a stable regime of training, and let the test error be the resulting…
- 0 votes0 replies0 views
Topological parallax and trustworthy deep-learning interpolation
The paper considers topological parallax as a comparison between a trained model and a reference dataset based on whether they have similar multiscale geometric structure. Topologi…
- 0 votes0 replies1 view
The equivalent-minima conjecture for large neural networks
Consider large neural networks trained by an optimization process, their local minima, and their generalization performance on a test set. Equivalent-minima conjecture. Large neura…
- 0 votes0 replies1 view
Stanislaw et al.'s learning-rate-to-batch-size conjecture for SGD
Let stochastic gradient descent (SGD) use a learning rate and a batch size, and consider the width of the endpoint reached by the optimization process and its generalization capabi…
- 0 votes0 replies1 view
Hochreiter's curvature conjecture for deep-network generalization
Let a deep neural network (a deepnet) be trained to a converged solution, and consider the curvature of its loss at that solution. Hochreiter et al.'s conjecture. The generalizatio…
- 0 votes0 replies0 views
The SGD flat-minima conjecture for generalization
A model is trained by a simple iterative method such as Stochastic Gradient Descent (SGD), and a flat local minimum is a local minimum associated with good generalization. Flat-min…
- 0 votes0 replies0 views
The implicit-bias conjecture for stochastic empirical risk minimization
Let an empirical risk minimization procedure use a stochastic learning algorithm, and call a minimizer generalizing when it performs well beyond the training data. Implicit-bias co…
- 0 votes0 replies0 views
The implicit-bias conjecture for gradient descent in deep learning
Overparameterized deep networks may have infinitely many global minimizers that fit the training samples exactly, so the optimization algorithm can determine which solution is obta…
- 0 votes0 replies0 views
The local-distance conjecture for autoencoder generalization
Let and be the training and test subsets of the dataset, and let be a trained autoencoder. Near each training…
- 0 votes0 replies0 views
The memorization explanation for degraded generalization at a skipped value
Memorization conjecture. The later deterioration is caused by the networks memorizing the answers for all the values except , thereby degrading performance at…
- 0 votes0 replies1 view
Conjectured improvement of the upper excess-risk bound for constant-stepsize SGD
Let be the number of SGD iterations, the constant stepsize, the initial iterate, the population-risk minimizer, and the data covariance matrix. Write…
- 0 votes0 replies0 views
Gradient Starvation's protection against overfitting
In the setting where training and test data points are drawn from the same distribution, the most salient features are predictive in both sets. Gradient Starvation's overfitting co…
- 0 votes0 replies0 views
Optimizer-induced implicit regularization in deep neural networks
A neural network estimator is obtained by optimizing an empirical loss over a class of deep neural networks. Optimizer-induced implicit regularization conjecture. It is conjectured…
- 0 votes0 replies0 views
The small-local-Lipschitz-constant search-path conjecture
Let a search path be a sequence of parameter values generated during optimization, and let its local Lipschitz constant measure the local smoothness along that path. Small-local-Li…
- 0 votes0 replies0 views
Conjecture that expected loss, error and generalization decrease with mini-batch size
Mini-batch performance monotonicity conjecture. There exists a decreasing property for the expected loss, error and the generalization ability with respect to the mini-batch size.