1 problem
- 0 votes0 replies0 views
Necessity of knowing the optimal bias span for span-dependent regret bounds
Let denote an upper bound on the span of the optimal bias function in an average-reward Markov decision process. In the online setting, regret is measured over the int…