3 problems
- 0 votes0 replies1 view
Boundedness of the proxy approximation between TD error and value gradients
Boundedness of the proxy approximation. There exists a constant such that, for a sufficiently fine state-space discretization,
- 0 votes0 replies0 views
Extension of the asymptotic bias formula to general eligibility traces
The TD learning recursion is governed by a Markov process, with eligibility-trace parameter ; for , the asymptotic bias limit exists and is given by the bias fo…
- 0 votes0 replies0 views
Extension of federated TD analysis to error-feedback encoding and realistic channels
The Quantized Federated TD learning algorithm, or QFedTD, updates the server parameter by … where is a constant step size, are independent Bernoulli packet-deliv…