Quick Best Action Identification In Linear Bandit Problems
2018 Β· Jun Geng, Lifeng Lai
Abstract
In this paper, we consider a best action identification problem in the stochastic linear bandit setup with a fixed confident constraint. In the considered best action identification problem, instead of minimizing the accumulative regret as done in existing works, the learner aims to obtain an accurate estimate of the underlying parameter based on his action and reward sequences. To improve the estimation efficiency, the learner is allowed to select his action based his historical information; hence the whole procedure is designed in a sequential adaptive manner. We first show that the existing algorithms designed to minimize the accumulative regret is not a consistent estimator and hence is not a good policy for our problem. We then characterize a lower bound on the estimation error for any policy. We further design a simple policy and show that the estimation error of the designed policy achieves the same scaling order as that of the derived lower bound.
Authors
(none)
Tags
Stats
Related papers
- Restless Bandit Problem With Rewards Generated By A Linear Gaussian Dynamical System (2024)0.00
- Optimal Policies For Observing Time Series And Related Restless Bandit Problems (2017)0.00
- Adaptive Sampling For Best Policy Identification In Markov Decision Processes (2020)0.00
- Towards Optimal Regret In Adversarial Linear Mdps With Bandit Feedback (2023)0.00
- Complete Policy Regret Bounds For Tallying Bandits (2022)0.00
- Multi-action Restless Bandits With Weakly Coupled Constraints: Simultaneous Learning And Control (2024)0.00
- Learning Near Optimal Policies With Low Inherent Bellman Error (2020)0.00
- Anti-concentrated Confidence Bonuses For Scalable Exploration (2021)0.00