109 problems
Let the linear multistep neural network loss have infinitely many global minimizers because the governing-function components are underdetermined by the training equations, a…
A CuBAS-selected warm-up set is a subset of the available training data, chosen using the CuBAS framework's curvature-based sampling procedure. Let the warm-up set comprise --…
A coefficient inverse problem (CIP) is an inverse problem for recovering coefficients of a differential equation from boundary measurements; here the approach combines a convexific…
Let , , and let or be continuous. Consider a Sprecher Network with architecture or…
Let the final-layer embeddings of a deep network be represented by their Gram matrix, and measure its complexity by effective rank. Huh et al. observe that this Gram matrix has low…
Let and let be an activation function with second Hermite coefficient . Exactness conjecture. For this one-hidden-layer perceptron, the replica symmetric th…
Let a deep-learning model be trained either with graphical processing units (GPUs) or with a central processing unit (CPU), and suppose the CPU-trained model incorporates sufficien…
U-NET optimal trade-off conjecture. A U-NET provides the optimal trade-off between performance and complexity for radio-to-PPG translation, especially given a limited-size dataset.
Loss decomposition conjecture. The loss may also be decomposed into two parts: one controlling the error from invariance, and the other controlling the error from the target.
Let denote the continuous piecewise linear functions on . The previously known upper bound is hidden layers.…
Redundant-encoding conjecture. More redundant encodings may be preferable for the network, as they could lead to representations that are more robust to small input variations and…
Balancedness is a structural property of deep linear networks, represented by the balanced manifold . Local equilibrium along a network is described as detailed balance.…
A deep linear network (DLN) is a neural network whose layers are linear maps, and deep learning refers to neural-network models trained on function-approximation tasks. Numerical e…
Let a neural operator have adjustable capacity, for example through its number of layers or neurons, and consider the non-convex optimization landscape of its training loss. Overpa…
In asynchronous pipeline training, let the parameter updates be subject to the delays induced by the pipeline execution. Delay-generalization conjecture. The observed acceleration…
Consider attention mechanisms in which data is exchanged between tokens. The authors entertain the attention-mechanism invariance conjecture. The choice of attention mechanism matt…
Lenses cannot model automatic differentiation algorithms that do not use gradient checkpointing, while weighted optics conjecture. weighted optics should be able to model such algo…
Let SG denote stochastic gradient and SGM denote stochastic gradient method with momentum. For deep neural networks (DNNs), consider objective functions, their ravines, the expecte…
Higher-order interactions (HOIs) are interactions involving more than two components in a complex system. Non-linear activation functions incorporating HOIs are related to attentio…
A hypernetwork is trained to predict optimized weights from hyperparameters, and its predictions are evaluated in regions of hyperparameter space with limited training data. Hypern…
The examples concern neural-network approximation of solution operators for parametric partial differential equations, including the parametric stationary Boussinesq equations, wit…
The absolute value function has derivative magnitude wherever it is differentiable. Gradient-explosion conjecture. Using the absolute value function in a training objective con…
Compositional embedding conjecture. A compositional framework to combine the parameter and response embeddings can lead to a more expressive model that can approximate more complex…
Lottery ticket hypothesis. Dense, randomly initialized, feed-forward networks contain winning-ticket subnetworks that, when trained in isolation, reach test accuracy comparable to…
Asymptotic covariance approximation conjecture. The population covariances can be asymptotically approximated by the last iterates of the corresponding linear recursions, with the…