90 problems
- 0 votes0 replies0 views
Minimax lower-bound conjecture for distributionally robust reinforcement learning
Minimax lower-bound conjecture. For distributionally robust reinforcement learning, the minimax lower bound on the number of samples is still
- 0 votes0 replies0 views
Jiang et al.'s horizon-dependent sample complexity conjecture for tabular reinforcement learning
In tabular reinforcement learning with planning horizon , consider any algorithm seeking an -optimal policy when the total reward is bounded by . Jiang et al.'s con…
- 0 votes0 replies1 view
Boundedness of the proxy approximation between TD error and value gradients
Boundedness of the proxy approximation. There exists a constant such that, for a sufficiently fine state-space discretization,
- 0 votes0 replies0 views
Low-curvature skill fixed-point conjecture for LASKO systems
Let be a Lie algebroid of controlled intervention modes over a typed Markdown workflow space, with anchor . For co…
- 0 votes0 replies0 views
Extension of the asymptotic bias formula to general eligibility traces
The TD learning recursion is governed by a Markov process, with eligibility-trace parameter ; for , the asymptotic bias limit exists and is given by the bias fo…
- 0 votes0 replies0 views
Optimal variance-dependent regret bound conjecture for infinite-horizon MDPs
Let and be integers, let , and consider horizon- algorithms for MDPs with states, actions, and diameter at mo…
- 0 votes0 replies1 view
The reward-hacking conjecture for Xent Games
A verifiable task is one in which a judge model can enforce the legality of player moves and score their plausibility, thereby providing a continuous reward signal for search optim…
- 0 votes0 replies1 view
Conjecture on the usefulness of ellipticity-induced Hilbert space structure
The paper studies model-free off-policy reinforcement learning in continuous time with general function approximation. Its analysis uses an ellipticity condition, which induces a H…
- 0 votes0 replies0 views
Conjecture on persistent bias in neural actor-critic algorithms
Persistent-bias conjecture. This behavior arises from the algorithm’s online learning nature and/or the coupled dynamics of the actor and critic networks, which induce a more persi…
- 0 votes0 replies0 views
Conjecture on gating rules from regularizer-dependent proximal gradients
Regularizer-dependent gating conjecture. Different gating rules will emerge from proximal gradients arising from different regularizers.
- 0 votes0 replies1 view
The Hamiltonian conjecture for continuous-time reinforcement learning
In continuous-time reinforcement learning, let the dynamics be described by a controlled diffusion and let the Hamiltonian denote the local operator combining the drift, diffusion,…
- 0 votes0 replies0 views
Unavoidability of minimum visitation probability dependence in corruption-tolerant asynchronous Q-learning
Let denote the minimum visitation probability in the asynchronous Q-learning setting described above. Visitation-probability dependence conjecture. Some de…
- 0 votes0 replies0 views
Extension of the analysis to general controlled Markov diffusions
Extension conjecture. The insights gained from this analysis could be extended to general controlled Markov diffusions beyond the fine-tuning setting.
- 0 votes0 replies1 view
Sequential empowerment inequality conjecture for the ICCEA power metric
Consider a sequential decision process, or Markov decision process, with state , action set , successor state , discount factor , and an empo…
- 0 votes0 replies1 view
Principal-solution conjecture for soft maximization in cyclic environments
Let and be the continuous self-maps defining the model-based and robot value equations, with positive inverse-temperature parameters and . Since and…
- 0 votes0 replies1 view
Stratified latent-space conjecture for reinforcement-learning state-action trajectories
Stratified latent-space conjecture. Distinct strata in the latent space correspond to different state-action trajectories, with increases in local dimension occurring when the agen…
- 0 votes0 replies0 views
The minimum-attention conjecture for meta-learning adaptation
Minimum attention is defined by the control-change functional … Here is a control varying over state space and time interval , and the functional penalizes chan…
- 0 votes0 replies0 views
Failure of the linear convergence rate under controlled volatility
Convergence-rate conjecture. The linear rate of convergence in the estimate referred to as, obtained for uncontrolled volatility, does not hold in the controlled-volatility case, f…
- 0 votes0 replies1 view
Conjecture that strong ergodicity can be weakened in risk-aware average-cost MDPs
Ergodicity conjecture. The strong ergodicity assumption could be weakened.
- 0 votes0 replies1 view
A smaller logarithmic factor may suffice in Zurek et al.'s concentration inequalities
Conjecture on improving . A smaller function may be sufficient for the cited inequalities to hold, so that the resulting improvement could be carried over to the p…
- 0 votes0 replies0 views
The wider-regime conjecture for minimax lower bounds in off-policy evaluation
Wider-regime conjecture. A similar lower bound should hold in the wider regime .
- 0 votes0 replies0 views
Optimal plug-in complexity for uniformly mixing MDPs
Uniformly mixing complexity conjecture. Theorem should also imply an optimal complexity of for this setting by an anal…
- 0 votes0 replies0 views
Maximum-likelihood optimal control outperforms direct policy learning in the Avellaneda–Stoikov model
The ergodic Avellaneda–Stoikov market-making model is assumed to describe market behaviour up to an unknown parameter, and the optimal control derived from maximum-likelihood estim…
- 0 votes0 replies0 views
The sampling-error explanation for nonmonotonic temperature effects
Let denote the temperature parameter controlling the weight on exploration, and consider the learned option prices obtained from the stopping and control procedures. For…
- 0 votes0 replies0 views
Instance-optimality of VRCQ for multiple actions
Let VRCQ denote the variance-reduced cascade Q-learning algorithm for estimating the optimal Q-function in a discounted Markov decision process, and let be its action…