29 problems
The TD learning recursion is governed by a Markov process, with eligibility-trace parameter ; for , the asymptotic bias limit exists and is given by the bias fo…
Let be convex and of class for an even integer , with unique minimizer , satisfying for…
Chen's conjecture. The established sufficient condition for strong consistency of the stochastic-gradient identification algorithm is inherently conservative; the…
Let be the moving boundary used by the textsc{StageSampling} procedure, which adaptively samples at a design point and stops when the accumulated statistic satisfie…
Let be the number of agents, let , and let be the vector of unreliability parameters. Let…
Bias characterization conjecture. The bias characterization result, including the conclusion that under non-degenerate noise the steady-state bias is of order…
The authors study temporal-difference learning with linear function approximation under Markovian sampling and develop an inductive technique for proving that the iterates remain u…
The paper studies inductive proofs for the effect of delays on iterative reinforcement-learning algorithms and compares them with robustness results for temporal-difference learnin…
A stochastic approximation problem has an underlying operator with some version of Lipschitzness when, for example, its operator is a gradient arising in smooth optimization. Adapt…
The Adam-family framework uses sequences of stepsizes and , and the sequence asymptotically approximates trajectories of a differential inclusi…
Let the techniques developed in this paper be applied to optimization problems with objectives having multiple local minima, and let simulated annealing be used as a comparison met…
Consider the stochastic approximation setting above, with convergence-rate parameters , , and satisfying . The established ra…
Average-utility convergence conjecture. For every such and ,
Sliding-mode implication conjecture. The existence of a sliding mode solution in the DI must imply that policy chattering takes place almost surely.
Invariant-chain-transitive-set conjecture. The only invariant, internally chain transitive sets of the DI $$ are the singleton sets
Let the stochastic approximation methods and the long-range dependent white noise processes considered in the paper be given, with convergence-rate exponent and Hurst inde…
Let denote the solution to the fixed-point equation arising in the preceding bound, and let its leading-order term refer to the dominant asymptotic contribution as become…
Boundary-utility conjecture. The following condition is sufficient: for all ,
Let and be the exponents governing the step-size and policy-update schedules, respectively. The stated convergence bound requires…
Consider the update rule … where and each is the probability of selecting state , with every state having no…
Risk-aware Q-learning (RaQL) uses an inner--outer loop structure in which the risk is estimated in an inner procedure and the -values are updated in an outer procedure. RaQL bia…
Risk-aware dynamic programming estimates risk at each iteration by solving an exact saddle-point optimization problem, while RaQL uses stochastic approximation with a stochastic ap…
Zap Q()-learning refers to the family of stochastic-approximation algorithms introduced in the paper for parameterized reinforcement-learning problems. The source asks for…
Let be the total number of queries required by the first tests in the modified probabilistic bisection algorithm (PBA), and let the error be measured by the distance betw…
Convergence to equilibrium. For any WARM with , there exists a random vector , supported on the set of linearly-stable and critical equilibria, s…