5 problems
- 0 votes0 replies0 views
Logarithmic ansatz conjecture for the optimal value function in mean-field control
Logarithmic ansatz conjecture. The optimal value function is conjectured to have the form
- 0 votes0 replies0 views
Unavoidability of minimum visitation probability dependence in corruption-tolerant asynchronous Q-learning
Let denote the minimum visitation probability in the asynchronous Q-learning setting described above. Visitation-probability dependence conjecture. Some de…
- 0 votes0 replies0 views
Instance-optimality of VRCQ for multiple actions
Let VRCQ denote the variance-reduced cascade Q-learning algorithm for estimating the optimal Q-function in a discounted Markov decision process, and let be its action…
- 0 votes0 replies0 views
Extension of constant-stepsize Q-learning results to linear function approximation
Constant-stepsize asynchronous Q-learning produces iterates whose distributional convergence, convergence-rate characterization, central limit theorem for averaged iterates, asympt…
- 0 votes0 replies1 view
RaQL's independent risk estimation and Q-value updates reduce iterative bias
Risk-aware Q-learning (RaQL) uses an inner--outer loop structure in which the risk is estimated in an inner procedure and the -values are updated in an outer procedure. RaQL bia…