Abstract:Over the past several decades, economists, psychologists, and neuroscientists have conducted experiments in which a subject, human or animal, repeatedly chooses between alternative actions and is rewarded based on choice history. While individual choices are unpredictable, aggregate behavior typically follows Herrnstein's matching law: the average reward per choice is equal for all chosen alternatives. In general, matching behavior does not maximize the overall reward delivered to the subject, and therefore matching appears inconsistent with the principle of utility maximization. Here we show that matching can be made consistent with maximization by regarding the choices of a single subject as being made by a sequence of multiple selves-one for each instant of time. If each self is blind to the state of the world and discounts future rewards completely, then the resulting game has at least one Nash equilibrium that satisfies both Herrnstein's matching law and the unpredictability of individual choices. This equilibrium is, in general, Pareto suboptimal, and can be understood as a mutual defection of the multiple selves in an intertemporal prisoner's dilemma. The mathematical assumptions about the multiple selves should not be interpreted literally as psychological assumptions. Human and animals do remember past choices and care about future rewards. However, they may be unable to comprehend or take into account the relationship between past and future. This can be made more explicit when a mechanism that converges on the equilibrium, such as reinforcement learning, is considered. Using specific examples, we show that there exist behaviors that satisfy the matching law but are not Nash equilibria. We expect that these behaviors will not be observed experimentally in animals and humans. If this is the case, the Nash equilibrium formulation can be regarded as a refinement of Herrnstein's matching law.

A dynamical policy search model for matching law.

A Stochastic Policy Search Model for Matching Behavior

Algorithm of matching law based on optimal policy search model

Matching provides efficient decisions

Operant matching as a Nash equilibrium of an intertemporal game

Matching-Based Policy Learning

Policy-Based Reinforcement Learning for Assortative Matching in Human Behavior Modeling

Towards Efficient Exact Optimization of Language Model Alignment

Reward Maximization in General Dynamic Matching Systems

An Approximate Dynamic Programming Approach to Dynamic Stochastic Matching

Policy Evaluation and Seeking for Multi-Agent Reinforcement Learning Via Best Response

Optimal bipartite graph matching-based goal selection for policy-based hindsight learning

A Policy Search Method For Temporal Logic Specified Reinforcement Learning Tasks

Achieving Correlated Equilibrium by Studying Opponent's Behavior Through Policy-Based Deep Reinforcement Learning

Reward Advancement: Transforming Policy under Maximum Causal Entropy Principle

Policy Learning for Balancing Short-Term and Long-Term Rewards

A two-sided matching decision-making approach based on prospect theory under the probabilistic linguistic environment

Online Matching with Stochastic Rewards: Advanced Analyses Using Configuration Linear Programs

Optimal Policy for Dynamic Assortment Planning under Multinomial Logit Models

Law Article-Enhanced Legal Case Matching: a Causal Learning Approach

A Model about Love in Human Dynamics