Lower-bound conjecture for adaptive linear-quadratic regulator regret

About 8 years old · traced to

Let π^\widehat{\pi} be an arbitrary adaptive policy, and let Rn(π^)\mathcal{R}_n(\widehat{\pi}) denote its regret after nn time points. Lower-bound conjecture. For an arbitrary adaptive policy π^\widehat{\pi} we have

lim inf⁡n→∞n−1/2Rn(π^)>0.\liminf\limits_{n \to \infty} n^{-1/2}\mathcal{R}_{n} \left(\widehat{\pi}\right)>0.

The conjecture asserts a universal lower bound on the regret growth of adaptive policies in linear-quadratic regulation, suggesting that no adaptive regulator can achieve regret of order smaller than n\sqrt{n}. The supplied text presents this as an interesting direction for future work and gives no resolution.

References

Primary source

Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari and George Michailidis, “On Adaptive Linear-Quadratic Regulators”, arXiv:1806.10749 (2020).

Progress summary

Never refreshed

Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.

Solutions 0

No solutions have been posted yet.