Abstract:Regret minimization has proved to be a versatile tool for tree-form sequential decision making and extensive-form games. In large two-player zero-sum imperfect-information games, modern extensions of counterfactual regret minimization (CFR) are currently the practical state of the art for computing a Nash equilibrium. Most regret-minimization algorithms for tree-form sequential decision making, including CFR, require (i) an exact model of the player's decision nodes, observation nodes, and how they are linked, and (ii) full knowledge, at all times t, about the payoffs -- even in parts of the decision space that are not encountered at time t. Recently, there has been growing interest towards relaxing some of those restrictions and making regret minimization applicable to settings for which reinforcement learning methods have traditionally been used -- for example, those in which only black-box access to the environment is available. We give the first, to our knowledge, regret-minimization algorithm that guarantees sublinear regret with high probability even when requirement (i) -- and thus also (ii) -- is dropped. We formalize an online learning setting in which the strategy space is not known to the agent and gets revealed incrementally whenever the agent encounters new decision points. We give an efficient algorithm that achieves $O(T^{3/4})$ regret with high probability for that setting, even when the agent faces an adversarial environment. Our experiments show it significantly outperforms the prior algorithms for the problem, which do not have such guarantees. It can be used in any application for which regret minimization is useful: approximating Nash equilibrium or quantal response equilibrium, approximating coarse correlated equilibrium in multi-player games, learning a best response, learning safe opponent exploitation, and online play against an unknown opponent/environment.

A note on continuous-time online learning

Online Mixed Discrete and Continuous Optimization: Algorithms, Regret Analysis and Applications

Online Bandit Learning against an Adaptive Adversary: from Regret to Policy Regret

Efficient Constrained Regret Minimization

Understanding the Role of Feedback in Online Learning with Switching Costs

Online Convex Optimization with Continuous Switching Constraint

Doubly Optimal No-Regret Online Learning in Strongly Monotone Games with Bandit Feedback

Online Learning with Feedback Graphs: Beyond Bandits

Continuous Prediction with Experts' Advice

Efficient Methods for Non-stationary Online Learning

Best-Case Lower Bounds in Online Learning

Model-Free Online Learning in Unknown Sequential Decision Making Problems and Games

Online Learning: Stochastic and Constrained Adversaries

Improving Adaptive Online Learning Using Refined Discretization

Online Control with Adversarial Disturbance for Continuous-time Linear Systems

Adaptive Regret for Bandits Made Possible: Two Queries Suffice

No-regret Learning for Repeated Non-Cooperative Games with Lossy Bandits

On Adaptivity in Information-constrained Online Learning

Constrained Online Two-stage Stochastic Optimization: Near Optimal Algorithms via Adversarial Learning

Online Stackelberg Optimization via Nonlinear Control

A Unified Framework for Analyzing Meta-algorithms in Online Convex Optimization