Data-dependent equivalence conjecture for random features

Let the assumptions in \Crefrange{asm:model}{asm:feat} hold. Define ϕn=p/n\phi_n=p/n. Let knk\le n be the subsample size and set ψˉn=p/k\bar{\psi}_n=p/k. Suppose that φ\varphi satisfies certain regularity conditions. For any MN{}M\in\mathbb{N}\cup\{\infty\}, define λˉn\bar{\lambda}_n by

1M=1M1ktr[(1kφ(LIXF)φ(LIXF))+]=1ntr[(1nφ(XF)φ(XF)+λˉnIn)1].\frac{1}{M}\sum_{\ell=1}^{M}\frac{1}{k}\operatorname{tr}\left[\left(\frac{1}{k}\varphi(\mathbf{L}_{I_{\ell}}\mathbf{X}\mathbf{F}^{\top})\varphi(\mathbf{L}_{I_{\ell}}\mathbf{X}\mathbf{F}^{\top})^{\top}\right)^+\right]=\frac{1}{n}\operatorname{tr}\left[\left(\frac{1}{n}\varphi(\mathbf{X}\mathbf{F}^{\top})\varphi(\mathbf{X}\mathbf{F}^{\top})^{\top}+\bar{\lambda}_n\mathbf{I}_n\right)^{-1}\right].

Define the data-dependent path Pn=P(λˉn;ϕn,ψˉn)\mathcal{P}_n=\mathcal{P}(\bar{\lambda}_n;\phi_n,\bar{\psi}_n). Random-feature equivalence conjecture. The conclusions of the risk-equivalence theorem, the corresponding prediction-risk proposition, and the linear-estimator equivalence theorem continue to hold for (φ(XF),y)(\varphi(\mathbf{X}\mathbf{F}^{\top}),\mathbf{y}), with P\mathcal{P} replaced by Pn\mathcal{P}_n. This conjecture extends the subsampling--ridge correspondence from the assumed random-matrix features to random features with nonlinear activation functions. Its status is not established in the supplied source.

Sources & referencesView supporting material

Primary source

Pratik Patil and Jin-Hong Du, “Generalized equivalences between subsampling and ridge regularization”, arXiv:2305.18496 (2023).

Progress summary

Never refreshed

Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.

Solutions 0

No solutions have been posted yet.