Convergence conjectures for deep linear neural network gradient flows
Convergence conjectures for deep linear neural network gradient flows
Assume that has full rank. Let be the depth, let denote the product along the gradient flow, and let denote the rank-constrained manifold used in Theorem 1.1. In the autoencoder case, assume , , let , and let be the eigenvalues of with corresponding orthonormal eigenvectors . Write for the matrix whose columns are . Convergence conjecture. (a) The statement in Theorem 1.2(b) also holds for . (b) If , has rank , and for every , then converges to . (c) In Theorem 1.3, if , convergence to a critical point on some with occurs only for a measure-zero set of initial values. These claims would sharpen the generic convergence theory for deep linear networks, extending the established depth-two result and describing the limiting autoencoder and lower-rank behavior. Their resolution is not supplied in the source.
Sources & referencesView supporting material
Primary source
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege and Michael Westdickenberg, “Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers”, arXiv:1910.05505 (2020).
Progress summary
Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.
Solutions 0
Sign in to submit a solution.
No solutions have been posted yet.