Feel-good thompson sampling for contextual bandits and reinforcement learning

Tong Zhang
2022-01-01
Abstract:Thompson sampling has been widely used for contextual bandit problems due to the flexibility of its modeling power. However, a general theory for this class of methods in the frequentist setting is still lacking. In this paper, we present a theoretical analysis of Thompson sampling, with a focus on frequentist regret bounds. In this setting, we show that the standard Thompson sampling is not aggressive enough in exploring new actions, leading to suboptimality in some pessimistic situations. A simple modification called Feel-Good Thompson sampling, which favors high reward models more aggressively than the standard Thompson sampling, is proposed to remedy this problem. We show that the theoretical framework can be used to derive Bayesian regret bounds for standard Thompson sampling and frequentist regret bounds for Feel-Good Thompson sampling. It is shown that in both cases, we can reduce the …
What problem does this paper attempt to address?