1 problem
- 0 votes0 replies0 views
Jiang et al.'s horizon-dependent sample complexity conjecture for tabular reinforcement learning
In tabular reinforcement learning with planning horizon , consider any algorithm seeking an -optimal policy when the total reward is bounded by . Jiang et al.'s con…