Covariance domination conjecture for the uniform prior over the ℓ1 ball

About 2 years old · traced to

Let CC be the ℓ1\mathbf{\ell}_1 ball, let p(w∣ξ)p(w\mid\xi) be the posterior density under the uniform prior on CC, and write Cov⁡Uni(C)(w)\operatorname{Cov}_{\mathrm{Uni}(C)}(w) for the covariance matrix of the uniform distribution on CC. For a direction vv, let z=v⋅wz=v\cdot w. Covariance domination conjecture. The posterior covariance is dominated by the prior covariance,

Cov⁡[w∣ξ]⪯Cov⁡Uni(C)(w).\operatorname{Cov}[w\mid\xi]\preceq \operatorname{Cov}_{\mathrm{Uni}(C)}(w).

Equivalently, for every direction vv,

Var⁡(v⋅w∣ξ)≤Var⁡Uni(C)(v⋅w).\operatorname{Var}(v\cdot w\mid\xi)\leq \operatorname{Var}_{\mathrm{Uni}(C)}(v\cdot w).

The authors present this as the uniform-prior analogue of their contraction result, but leave it for future work because the geometry of the convex constraint makes the one-dimensional marginals and their Hessians more complicated.

References

Primary source

Curtis McDonald and Andrew R Barron, “Log-Concave Coupling for Sampling Neural Net Posteriors”, arXiv:2407.18802 (2024).

Progress summary

Refreshed
Claimed solved

A 2026 posted calculation claims the conjecture is false in every dimension from two upward, but no independent verification was found.

McDonald and Barron introduced this covariance-domination conjecture for the uniform prior on the unit ℓ1\ell_1 ball as a future problem in their 2024 paper. It asserts that conditioning on the auxiliary data cannot increase posterior variance in any direction.

Posted attempt

A reader-written calculation claims a complete counterexample: for every dimension d≥2d\ge2, a permitted one-observation construction makes a transverse posterior variance strictly exceed its uniform-prior value. It also claims the analogous auxiliary assertion in the authors’ 2026 follow-up fails under all its stated hypotheses, but the argument has not been independently verified.

Current status (as of August 2026): The conjecture has no verified resolution; an unverified posted calculation claims a counterexample in every dimension d≥2d\ge2, leaving the mathematical claim unsettled.

Sources

Solutions 1

CounterexampleThis solution needs a summarySee full solutionHide full solution

Counterexample in every dimension. Let Cd={w∈Rd:∥w∥1≤1}C_d=\{w\in\mathbb R^d:\|w\|_1\le1\}, and let WW be uniform on CdC_d. For U=∣W1∣U=|W_1|, the exact marginal density and transverse conditional second moments are

fU(u)=d(1−u)d−1,E[Wj2∣U=u]=2(1−u)2d(d+1)(0<u<1, j≥2).f_U(u)=d(1-u)^{d-1},\qquad \mathbb E[W_j^2\mid U=u]=\frac{2(1-u)^2}{d(d+1)} \quad(0<u<1,\ j\ge2).

In particular,

Var⁡Unif(Cd)(Wj)=2(d+1)(d+2).\operatorname{Var}_{\mathrm{Unif}(C_d)}(W_j) =\frac{2}{(d+1)(d+2)}.

In Conjecture 1 of the original source, take any d≥2d\ge2, one observation, and its allowed parameters

x1=e1,r1=−1,0<α<1,ψ(t)=t2/2,c=1,ξ1=0.x_1=e_1,\quad r_1=-1,\quad 0<\alpha<1,\quad \psi(t)=t^2/2,\quad c=1,\quad \xi_1=0.

Direct substitution into source equations (6)–(7) gives the reverse conditional density

p(w∣0)∝e−αw121{∥w∥1≤1}.p(w\mid0)\propto e^{-\alpha w_1^2}\mathbf1_{\{\|w\|_1\le1\}}.

Both u↦(1−u)2u\mapsto(1-u)^2 and u↦e−αu2u\mapsto e^{-\alpha u^2} strictly decrease on (0,1)(0,1). The strict covariance identity

Cov⁡(f(U),g(U))=12E[(f(U)−f(U′))(g(U)−g(U′))]>0\operatorname{Cov}(f(U),g(U)) =\frac12\mathbb E[(f(U)-f(U'))(g(U)-g(U'))]>0

for an independent copy U′U' therefore gives, for every j≥2j\ge2,

Var⁡(Wj∣0)=2d(d+1)E[(1−U)2e−αU2]E[e−αU2]>2(d+1)(d+2)=Var⁡Unif(Cd)(Wj).\operatorname{Var}(W_j\mid0) =\frac{2}{d(d+1)} \frac{\mathbb E[(1-U)^2e^{-\alpha U^2}]} {\mathbb E[e^{-\alpha U^2}]} > \frac{2}{(d+1)(d+2)} =\operatorname{Var}_{\mathrm{Unif}(C_d)}(W_j).

Thus the asserted covariance domination fails strictly in every dimension d≥2d\ge2. All stated source restrictions hold, including d>αcn∥r∥∞d>\alpha cn\|r\|_\infty. For d=2,α=1/2d=2,\alpha=1/2, an exact alternating-series certificate even yields

Var⁡(W2∣0)≥16+7310080.\operatorname{Var}(W_2\mid0)\ge\frac16+\frac{73}{10080}.

The failure persists with a full-rank design and a strictly concave conditional log density: taking two observations x1=e1,x2=e2x_1=e_1,x_2=e_2, residuals −1,−1/100-1,-1/100, and the same α,ψ,c\alpha,\psi,c gives density proportional to

exp⁡(−w12/2−w22/200)1C2(w),\exp(-w_1^2/2-w_2^2/200)\mathbf1_{C_2}(w),

with transverse variance at least 1/6+12847/2016000>1/61/6+12847/2016000>1/6.

The latest published, truncated version also fails under its full hypotheses. The authors retain the analogous auxiliary assertion as equation (4.44) of their 2026 published follow-up, but its reverse conditional now has an additional truncation-normalization factor. To account for that factor and satisfy every hypothesis of its Theorem 1, choose

K=2, N=1, d=100000, δ=1/20,β=2log⁡(8000000),K=2,\ N=1,\ d=100000,\ \delta=1/20,\quad \beta=2\log(8000000), x1=e1,y1=0,ψ(t)=t,V=a0=a1=a2=1,c1=c2=1/2.x_1=e_1,\quad y_1=0,\quad \psi(t)=t,\quad V=a_0=a_1=a_2=1,\quad c_1=c_2=1/2.

Then CN=1C_N=1, A3=96eA_3=96\sqrt e, and ρ=β\rho=\beta. The required size restrictions hold:

K[log⁡(2Kd)+log⁡(1/δ)]=βN,K[\log(2Kd)+\log(1/\delta)]=\beta N, A3(βN)2<96⋅2⋅322=196608<200000=Kd,A_3(\beta N)^2 <96\cdot2\cdot32^2 =196608<200000=Kd,

and βN>2, K=2, H(1/20)<1/2\beta N>2,\ K=2,\ H(1/20)<1/2.

The truncation region is B=[−L,L]2B=[-L,L]^2 for some L>0L>0. For Tu∼N(u,ρ−1)T_u\sim N(u,\rho^{-1}), put

P(u)=Pr⁡(∣Tu∣≤L),s(u)=e−ρu2/2P(u).P(u)=\Pr(|T_u|\le L), \qquad s(u)=\frac{e^{-\rho u^2/2}}{P(u)}.

Differentiating the Gaussian integral gives

(log⁡s)′(u)=−ρ E[Tu∣∣Tu∣≤L]<0(u>0);(\log s)'(u) =-\rho\,\mathbb E[T_u\mid |T_u|\le L]<0 \quad(u>0);

the final strict sign follows by pairing tt with −t-t. Writing ui=wi,1u_i=w_{i,1}, the joint conditional density after integrating transverse coordinates is proportional to

h(u1)h(u2)e−β(u1+u2)2/8,h(u)=(1−∣u∣)d−1s(u)1{∣u∣<1}.h(u_1)h(u_2)e^{-\beta(u_1+u_2)^2/8}, \qquad h(u)=(1-|u|)^{d-1}s(u)\mathbf1_{\{|u|<1\}}.

Hence the first neuron's marginal, relative to its uniform prior, is weighted by

F(u)=s(u)(h∗gβ)(u),gβ(t)=e−βt2/8.F(u)=s(u)(h*g_\beta)(u), \qquad g_\beta(t)=e^{-\beta t^2/8}.

Convolution preserves the class of even functions decreasing on [0,∞)[0,\infty), as follows immediately from the layer-cake representation by centered intervals. Since ss is strictly decreasing there, so is FF. Applying the same strict slice-covariance argument above shows

Var⁡p∗( ⋅ ∣0)(w1,j)>Var⁡Unif(Cd)(w1,j)(j≥2).\operatorname{Var}_{p^*(\,\cdot\,\mid0)}(w_{1,j}) > \operatorname{Var}_{\mathrm{Unif}(C_d)}(w_{1,j}) \quad(j\ge2).

Consequently, the latest published equation (4.44) fails even under every stated size and truncation hypothesis.

A separate variance error. Both papers print d/((d+1)2(d+2))d/((d+1)^2(d+2)) for a signed uniform coordinate. That expression is instead Var⁡(∣Wj∣)\operatorname{Var}(|W_j|); the correct signed variance is 2/((d+1)(d+2))2/((d+1)(d+2)), also implied by the published paper's own Appendix equation (7.54).

These counterexamples refute the auxiliary covariance-domination assertions, not the main theorem of the later paper, which uses a different bound.

Sources: original Conjecture 1; final published article, equation (4.44).