Gaussian equivalence conjecture for the empirical conjugate kernel matrix

From papers

Let X~\tilde{\boldsymbol X} and y~\tilde{\boldsymbol y} be new training data and labels independent of W1\boldsymbol W_1. Assume the activation σ\sigma is odd, Assumption 1 holds, and [?][?] the learning rate satisfies [?][?] [?][?]; define the empirical feature matrices

Φ=1Nσ(X~W1),Φ=1N\originalleft(μ1X~W1+μ2Z\aftergroup\originalright),\boldsymbol\Phi=\frac{1}{\sqrt N}\sigma(\tilde{\boldsymbol X}\boldsymbol W_1),\qquad \overline{\boldsymbol\Phi}=\frac{1}{\sqrt N}\mathopen{}\mathclose\bgroup\originalleft(\mu_1\tilde{\boldsymbol X}\boldsymbol W_1+\mu_2\boldsymbol Z\aftergroup\egroup\originalright),

where the entries of Z\boldsymbol Z are independent standard Gaussian variables. Let u1\boldsymbol u_1 and u1\overline{\boldsymbol u}_1 be the respective left leading singular vectors. Gaussian equivalence conjecture. The empirical feature matrix and its Gaussian equivalent have asymptotically matching singular values and leading-vector alignment with the labels:

\originalleftsi(Φ)si(Φ)\aftergroup\originalright=od,P(1)for all i[n],\mathopen{}\mathclose\bgroup\originalleft|s_i(\boldsymbol\Phi)-s_i(\overline{\boldsymbol\Phi})\aftergroup\egroup\originalright|=o_{d,\mathbb P}(1)\quad\text{for all }i\in[n],

and

\originalleft\originalleftu1,y~y~\aftergroup\originalright\aftergroup\originalright2=\originalleft\originalleftu1,y~y~\aftergroup\originalright\aftergroup\originalright2+od,P(1).\mathopen{}\mathclose\bgroup\originalleft|\mathopen{}\mathclose\bgroup\originalleft\langle\boldsymbol u_1,\frac{\tilde{\boldsymbol y}}{\|\tilde{\boldsymbol y}\|}\aftergroup\egroup\originalright\rangle\aftergroup\egroup\originalright|^2=\mathopen{}\mathclose\bgroup\originalleft|\mathopen{}\mathclose\bgroup\originalleft\langle\overline{\boldsymbol u}_1,\frac{\tilde{\boldsymbol y}}{\|\tilde{\boldsymbol y}\|}\aftergroup\egroup\originalright\rangle\aftergroup\egroup\originalright|^2+o_{d,\mathbb P}(1).

This conjecture proposes that Gaussian equivalence precisely captures the BBP-type spectral transition of the empirical conjugate kernel matrix when the population covariance contains a spike. It is motivated by the corresponding Gaussian-equivalent feature model and the preceding population-level transition; its general validity remains open.

Progress summary

Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.

Sources & referencesView supporting material

Primary source

Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu and Greg Yang, “High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation”, arXiv:2205.01445 (2022).

Solutions 0

No solutions have been posted yet.