Regret bound for the adaptive online policy in Algorithm 5
Regret bound for the adaptive online policy in Algorithm 5
Let be an independent and identically distributed process, and let follow either a linear regression model with white noise or a weighted random-walk model. Let denote the online policy specified by Algorithm 5, and let be its regret over rounds. Under suitable regularity conditions, Regret-bound conjecture.
The conjecture proposes a formal regret guarantee for the adaptive algorithm under the stochastic input models considered in the paper; the specific regularity conditions and a proof remain to be established.
Sources & referencesView supporting material
Primary source
Owen Shen, “Online Regenerative Learning”, arXiv:2209.08657 (2022).
Progress summary
Nothing recorded yet. Refresh searches the literature and the public web for attempts on this problem, and writes the first summary here.
Solutions 0
Sign in to submit a solution.
No solutions have been posted yet.