3 problems
- 0 votes0 replies1 view
Conjecture that strong ergodicity can be weakened in risk-aware average-cost MDPs
Ergodicity conjecture. The strong ergodicity assumption could be weakened.
- 0 votes0 replies1 view
RaQL's independent risk estimation and Q-value updates reduce iterative bias
Risk-aware Q-learning (RaQL) uses an inner--outer loop structure in which the risk is estimated in an inner procedure and the -values are updated in an outer procedure. RaQL bia…
- 0 votes0 replies0 views
The computational advantage of RaQL over exact risk-aware dynamic programming
Risk-aware dynamic programming estimates risk at each iteration by solving an exact saddle-point optimization problem, while RaQL uses stochastic approximation with a stochastic ap…