Weak identifiability of multi-head attention parameters
Weak identifiability of multi-head attention parameters
Let be the parameter space of a multi-head attention (MHA) model with value dimension and heads. Let be the permutation group on the heads, and let
A parameter is weakly identifiable with respect to when it is weakly identifiable after accounting for this symmetry group.
MHA weak identifiability conjecture. For every parameter of an MHA model, the parameter is weakly identifiable with respect to the symmetry group
The conjecture asserts that the parameterization map is well behaved after quotienting by the value-output general linear transformations and head permutations. The source presents identifiability for attention mechanisms as a nascent field and motivates omitting query and key symmetries because equivariance is typically mediated through the value-output path or positional encoding.
Sources & referencesView supporting material
Primary source
Vahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman and Kathlén Kohn, “Identifiable Equivariant Networks are Layerwise Equivariant”, arXiv:2601.21645 (2026).
Progress summary
Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.
Solutions 0
Sign in to submit a solution.
No solutions have been posted yet.