Abstract:We consider realizable contextual bandits with general function approximation, investigating how small reward variance can lead to better-than-minimax regret bounds. Unlike in minimax bounds, we show that the eluder dimension $d_\text{elu}$$-$a complexity measure of the function class$-$plays a crucial role in variance-dependent bounds. We consider two types of adversary: (1) Weak adversary: The adversary sets the reward variance before observing the learner's action. In this setting, we prove that a regret of $\Omega(\sqrt{\min\{A,d_\text{elu}\}\Lambda}+d_\text{elu})$ is unavoidable when $d_{\text{elu}}\leq\sqrt{AT}$, where $A$ is the number of actions, $T$ is the total number of rounds, and $\Lambda$ is the total variance over $T$ rounds. For the $A\leq d_\text{elu}$ regime, we derive a nearly matching upper bound $\tilde{O}(\sqrt{A\Lambda}+d_\text{elu})$ for the special case where the variance is revealed at the beginning of each round. (2) Strong adversary: The adversary sets the reward variance after observing the learner's action. We show that a regret of $\Omega(\sqrt{d_\text{elu}\Lambda}+d_\text{elu})$ is unavoidable when $\sqrt{d_\text{elu}\Lambda}+d_\text{elu}\leq\sqrt{AT}$. In this setting, we provide an upper bound of order $\tilde{O}(d_\text{elu}\sqrt{\Lambda}+d_\text{elu})$. Furthermore, we examine the setting where the function class additionally provides distributional information of the reward, as studied by Wang et al. (2024). We demonstrate that the regret bound $\tilde{O}(\sqrt{d_\text{elu}\Lambda}+d_\text{elu})$ established in their work is unimprovable when $\sqrt{d_{\text{elu}}\Lambda}+d_\text{elu}\leq\sqrt{AT}$. However, with a slightly different definition of the total variance and with the assumption that the reward follows a Gaussian distribution, one can achieve a regret of $\tilde{O}(\sqrt{A\Lambda}+d_\text{elu})$.

Variance-Aware Sparse Linear Bandits.

How Does Variance Shape the Regret in Contextual Bandits?

Variance-Dependent Regret Bounds for Non-stationary Linear Bandits

Sparsity-Agnostic Linear Bandits with Adaptive Adversaries

Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed Payoffs

Variance-Aware Confidence Set: Variance-Dependent Bound for Linear Bandits and Horizon-Free Bound for Linear Mixture MDP

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

Low-Rank Generalized Linear Bandit Problems

Nash Regret Guarantees for Linear Bandits

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits

Information Directed Sampling for Sparse Linear Bandits

Nearly Minimax-Optimal Regret for Linearly Parameterized Bandits.

Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery

Regret Minimization and Statistical Inference in Online Decision Making with High-dimensional Covariates

Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments

Geometry-Aware Approaches for Balancing Performance and Theoretical Guarantees in Linear Bandits

Linear bandits with polylogarithmic minimax regret

Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards

Risk-averse Contextual Multi-armed Bandit Problem with Linear Payoffs

Zero-Inflated Bandits