3 problems
- 0 votes0 replies0 views
Linear convergence conjecture for RMSProp and Adam on the softmax policy gradient objective
Let the softmax policy gradient objective be the objective considered in the paper, and let and denote the corresponding optimization methods. RM…
- 0 votes0 replies0 views
Conjecture on the noise sensitivity of policy-gradient estimates
Noise-sensitivity conjecture. The gradient estimates are more sensitive to noise than the discounted rewards.
- 0 votes0 replies0 views
Conjecture on the bounded monotonicity of policy-gradient oracle variance
Variance monotonicity conjecture. Under proper assumptions, the variance of the policy-gradient oracle is an increasing but bounded function of .