10 problems
- 0 votes0 replies0 views
The necessity of forgetting for optimal globally private bandit algorithms
In an -global differentially private stochastic bandit, an algorithm is said to forget when it discards past rewards between independent episodes. Forgetting-necessity con…
- 0 votes0 replies0 views
Optimal-regret conjecture for the TP policy in dynamic matching
In a two-way dynamic matching network, let denote the local availability-based policy proposed by Kerimov et al. The policy makes matching decisions using agent avail…
- 0 votes0 replies0 views
Dai et al.'s minimax Neyman regret rate conjecture
Let denote the number of units in a finite-population adaptive randomized controlled trial, and let denote a regret rate up to polylogarithmic factors…
- 0 votes0 replies0 views
Relative-regret extension conjecture for the censored newsvendor problem
Relative-regret extension conjecture. The paper's techniques should still apply to the relative regret metric.
- 0 votes0 replies0 views
Maximum-likelihood optimal control outperforms direct policy learning in the Avellaneda–Stoikov model
The ergodic Avellaneda–Stoikov market-making model is assumed to describe market behaviour up to an unknown parameter, and the optimal control derived from maximum-likelihood estim…
- 0 votes0 replies0 views
The aggregation precision conjecture for deteriorating Markov decision processes
Let denote the number of observed trajectories, and let and be, respectively, highly aggregated and more finely aggregated models of the same deteriora…
- 0 votes0 replies0 views
Arlotto et al.'s logarithmic-regret conjecture for the finite-type multi-secretary problem
Let be the time horizon. In the finite-type multi-secretary problem, candidates have abilities drawn from a distribution with finitely many types, and regret measures the loss…
- 0 votes0 replies0 views
Asymptotic optimality of the gradient feedback strategy
Gradient-strategy conjecture.
- 0 votes0 replies0 views
Asymptotic optimality of Hamiltonian-maximizing nature strategies
Let be the value function, let denote the admissible expert subsets, and define the set of Hamiltonian maximizers by … For and stopping parameter…
- 0 votes0 replies1 view
Worst-case suboptimality factor for ETC strategies
Let -armed bandit means satisfy and suppose that and are much larger than the other means. Worst-case ETC conjecture. In this regime, the regret of an…