8 problems
- 0 votes0 replies0 views
Allocation variability and slow CLT convergence for UCB1
UCB1 allocation-variability conjecture. The large allocation variability of texttt{UCB1} contributes to its slow convergence to the central limit theorem, and the convergence rate…
- 0 votes0 replies0 views
Pareto-efficient trade-offs for modified bandit algorithms
Conjecture on modified algorithms. Suitably modified variants of bandit algorithms should similarly achieve Pareto-efficient trade-offs as those achieved by texttt{UCB-f}.
- 0 votes0 replies0 views
Optimality conjecture for RS-DR and RS-SA in best-arm identification
Consider the proposed RS-AIPW strategy for fixed-budget best-arm identification in multi-armed bandits, together with the RS-DR and RS-SA strategies. RS-DR and RS-SA asymptotic opt…
- 0 votes0 replies0 views
RS-DR performance conjecture for re-estimated allocation probabilities
The setting is a two-armed bandit problem using the RS-DR strategy, with allocation probability for arm at time and a re-estimated allocation probability used by…
- 0 votes0 replies0 views
Eluder dimension requires sublinear information gain for sublinear regret
Eluder-dimension conjecture. The eluder dimension does not yield sublinear regret unless the information gain is sublinear in .
- 0 votes0 replies0 views
Phase-transition conjecture for optimal regret in multi-stage DTR bandits
Let denote the number of stages in a multi-stage dynamic treatment regime (DTR) bandit problem, and let denote the time horizon. As becomes very large, for example comp…
- 0 votes0 replies1 view
Conjecture on the tightness of the online influence maximization regret bound
Tightness conjecture. This regret bound is at most away from being tight.
- 0 votes0 replies0 views
The KL-HHR performance conjecture for stochastic online shortest path routing
KL-HHR performance conjecture. The performance of the algorithm should be very close to that of the algorithm, as observed in numerical exper…