Abstract:We consider realizable contextual bandits with general function approximation, investigating how small reward variance can lead to better-than-minimax regret bounds. Unlike in minimax bounds, we show that the eluder dimension $d_\text{elu}$$-$a complexity measure of the function class$-$plays a crucial role in variance-dependent bounds. We consider two types of adversary: (1) Weak adversary: The adversary sets the reward variance before observing the learner's action. In this setting, we prove that a regret of $\Omega(\sqrt{\min\{A,d_\text{elu}\}\Lambda}+d_\text{elu})$ is unavoidable when $d_{\text{elu}}\leq\sqrt{AT}$, where $A$ is the number of actions, $T$ is the total number of rounds, and $\Lambda$ is the total variance over $T$ rounds. For the $A\leq d_\text{elu}$ regime, we derive a nearly matching upper bound $\tilde{O}(\sqrt{A\Lambda}+d_\text{elu})$ for the special case where the variance is revealed at the beginning of each round. (2) Strong adversary: The adversary sets the reward variance after observing the learner's action. We show that a regret of $\Omega(\sqrt{d_\text{elu}\Lambda}+d_\text{elu})$ is unavoidable when $\sqrt{d_\text{elu}\Lambda}+d_\text{elu}\leq\sqrt{AT}$. In this setting, we provide an upper bound of order $\tilde{O}(d_\text{elu}\sqrt{\Lambda}+d_\text{elu})$. Furthermore, we examine the setting where the function class additionally provides distributional information of the reward, as studied by Wang et al. (2024). We demonstrate that the regret bound $\tilde{O}(\sqrt{d_\text{elu}\Lambda}+d_\text{elu})$ established in their work is unimprovable when $\sqrt{d_{\text{elu}}\Lambda}+d_\text{elu}\leq\sqrt{AT}$. However, with a slightly different definition of the total variance and with the assumption that the reward follows a Gaussian distribution, one can achieve a regret of $\tilde{O}(\sqrt{A\Lambda}+d_\text{elu})$.

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

Variance-Aware Confidence Set: Variance-Dependent Bound for Linear Bandits and Horizon-Free Bound for Linear Mixture MDP

Improved Algorithms for Stochastic Linear Bandits Using Tail Bounds for Martingale Mixtures

Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments

Variance-Aware Sparse Linear Bandits.

Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition

How Does Variance Shape the Regret in Contextual Bandits?

Variance-Dependent Regret Bounds for Non-stationary Linear Bandits

Horizon-Free and Variance-Dependent Reinforcement Learning for Latent Markov Decision Processes

Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits

Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization

Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion

Nearly Minimax-Optimal Regret for Linearly Parameterized Bandits.

Nearly Minimax Optimal Reinforcement Learning for Linear Mixture Markov Decision Processes

Data-Driven Upper Confidence Bounds with Near-Optimal Regret for Heavy-Tailed Bandits

Thompson Sampling Algorithms for Mean-Variance Bandits

Continuous Mean-Covariance Bandits.

Improving Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms and Its Applications

Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span

Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits