Existence of double scaling limit for random neural networks

About 4 years old · traced to

Consider a random depth LL neural network with input dimension n0n_0, hidden layer widths

n1,…,nL=n≫1,n_1,\ldots,n_L=n\gg 1,

output dimension nL+1n_{L+1} and non-linearity σ\sigma. Suppose that the network is tuned to criticality, meaning that the criticality condition is satisfied. Fix a non-zero network input xα∈Rn0x_\alpha\in\mathbb R^{n_0} and write ξ=L/n\xi=L/n. For each k≥1k\geq 1, let κ2k;α(L)\kappa_{2k;\alpha}^{(L)} denote the corresponding even cumulant and let K2;α(L)K_{2;\alpha}^{(L)} denote the corresponding second moment. Existence of the double scaling limit. For each k≥1k\geq 1 there exists C2k>0C_{2k}>0, depending on the universality class of σ\sigma, such that

κ2k;α(L)(K2;α(L))k=C2kξk−1+O(ξk).\frac{\kappa_{2k;\alpha}^{(L)}}{\left(K_{2;\alpha}^{(L)}\right)^k}=C_{2k}\xi^{k-1}+O\left(\xi^k\right).

Moreover, for each ξ∈[0,∞)\xi\in[0,\infty) there exists a probability distribution Pξ,σ\mathbb P_{\xi,\sigma} on R\mathbb R, depending only on ξ\xi and σ\sigma, such that in the double scaling limit

n,L⟶∞,Ln⟶ξ,n,L\longrightarrow\infty,\qquad \frac{L}{n}\longrightarrow\xi,

the random variable zi;α(ℓ)z_{i;\alpha}^{(\ell)} converges in distribution to a random variable with law Pξ,σ\mathbb P_{\xi,\sigma}. This conjecture proposes a universal limiting description of critically tuned random neural networks when width and depth grow proportionally; the existence of the limiting distribution and the stated cumulant asymptotics are not established in the supplied text.

References

Primary source

Boris Hanin, “Random Fully Connected Neural Networks as Perturbatively Solvable Hierarchies”, arXiv:2204.01058 (2023).

Progress summary

Never refreshed

Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.

Solutions 0

No solutions have been posted yet.