11 problems
Let be the teacher model's coefficient vector, with sampled in the paper's model from . Worst-case fine-tuning tractability conjecture. Fine-tuning sh…
While the analysis concerns analytically tractable overparameterized linear regression, let signal-noise dynamics denote the interaction between denoising and signal forgetting dur…
Let denote sieve spaces used for boosting, and let the associated hyperparameters be selected from the data, for example by sample-splitting or cross-validation. Da…
Let be an activation function satisfying the monomial approximation property: for every integer and every , there is a bounded weight function…
Let be the limiting parameter distribution at time , and suppose that the origin is an interior point of…
Let the data have modes satisfying the additional homogeneity assumption described in the paper, and let the training objective use a loss function that is not necessarily convex.…
Signal-strength tightness conjecture. The second condition on the signal strength is tight up to an extra factor; this factor is believed to be an artifact of the analys…
Consider the polynomial approximation barrier in high-dimensional kernel regression, and allow different choices of the kernel scaling, including the standard scaling and the flat-…
Let a kernel be rotationally invariant, and consider kernel regression for challenging non-polynomial problems in high dimensions. The polynomial approximation barrier is a bias ph…
Let the counterexample constructed in the paper establish failure of initial-value stability for standard deterministic accelerated gradient descent, even for strongly convex and s…
Let a base learner predict by randomly selecting and averaging over points in a set , and let denote the number of selected points. Write and for the cor…