Generic identifiability of normalized deep self-attention parametrizations

Let normalized deep self-attention networks be parametrized via (M,L)(\mathbf{M},L), where these parameters encode the network as defined in the paper. The normalization removes layer-wise scaling, while the remaining symmetries inside the attention matrices and the transformations CiC_i persist. Generic identifiability conjecture. For normalized deep self-attention networks, the generic fibers of the parametrization via (M,L)(\mathbf{M},L) are singletons; equivalently, the parametrization is generically one-to-one. This asserts that normalization only breaks the layer-wise scaling symmetry of the parametrization, extending the single-layer identifiability result to the deep case.

Sources & referencesView supporting material

Primary source

Nathan W. Henry, Giovanni Luca Marchetti and Kathlén Kohn, “Geometry of Lightning Self-Attention: Identifiability and Dimension”, arXiv:2408.17221 (2026).

Progress summary

Never refreshed

Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.

Solutions 0

No solutions have been posted yet.