2 problems
- 0 votes0 replies0 views
The wider-regime conjecture for minimax lower bounds in off-policy evaluation
Wider-regime conjecture. A similar lower bound should hold in the wider regime .
- 0 votes0 replies1 view
Robustness of an -based Bellman-error estimator under heavy-tailed rewards
Let be a -function and let be the reward function. An estimator can be constructed by minimizing an empirical approximation of the -norm of the Bellman error (resid…