Generic identifiability of normalized deep self-attention parametrizations
Generic identifiability of normalized deep self-attention parametrizations
Let normalized deep self-attention networks be parametrized via , where these parameters encode the network as defined in the paper. The normalization removes layer-wise scaling, while the remaining symmetries inside the attention matrices and the transformations persist. Generic identifiability conjecture. For normalized deep self-attention networks, the generic fibers of the parametrization via are singletons; equivalently, the parametrization is generically one-to-one. This asserts that normalization only breaks the layer-wise scaling symmetry of the parametrization, extending the single-layer identifiability result to the deep case.
Sources & referencesView supporting material
Primary source
Nathan W. Henry, Giovanni Luca Marchetti and Kathlén Kohn, “Geometry of Lightning Self-Attention: Identifiability and Dimension”, arXiv:2408.17221 (2026).
Progress summary
Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.
Solutions 0
Sign in to submit a solution.
No solutions have been posted yet.