6 problems
- 0 votes0 replies0 views
Optimal regret with time-varying perturbations in contextual dynamic pricing
Time-varying perturbation conjecture. Optimal regret could still be achieved using
- 0 votes0 replies0 views
Simulation-based confidence intervals after semiparametric adaptive stopping
The procedure discussed above concerns inference after adaptive data collection in a contextual bandit with stopping time : one simulates the limiting stopping-time distribut…
- 0 votes0 replies0 views
Thompson-sampling conjecture for contextual dueling bandits
Thompson-sampling conjecture. A Thompson-sampling-based algorithm may also work for the contextual dueling bandit setting because its stochastic nature encourages the agent to expl…
- 0 votes0 replies0 views
Semiparametric efficient score conjecture for contextual best-arm identification
Let and . Let denote the loss for arm at time horizon , and let the semiparametric efficient score functio…
- 0 votes0 replies0 views
Conjecture on singular price-covariate data causing RMLP-2 instability
Let RMLP-2 be a contextual dynamic pricing method that assumes a log-concave noise distribution, and consider the simulation settings with standard Gaussian noise in…
- 0 votes0 replies0 views
Conjecture that the distribution-free dynamic pricing regret upper bound is near-optimal
Regret lower-bound conjecture. The obtained regret upper bound is close to the lower bound for this setting. The problem is harder than standard linear bandits and dynamic pricing…