Abstract:Learned communication between agents is a powerful tool when approaching decision-making problems that are hard to overcome by any single agent in isolation. However, continual coordination and communication learning between machine agents or human-machine partnerships remains a challenging open problem. As a stepping stone toward solving the continual communication learning problem, in this paper we contribute a multi-faceted study into what we term Pavlovian signalling -- a process by which learned, temporally extended predictions made by one agent inform decision-making by another agent with different perceptual access to their shared environment. We seek to establish how different temporal processes and representational choices impact Pavlovian signalling between learning agents. To do so, we introduce a partially observable decision-making domain we call the Frost Hollow. In this domain a prediction learning agent and a reinforcement learning agent are coupled into a two-part decision-making system that seeks to acquire sparse reward while avoiding time-conditional hazards. We evaluate two domain variations: 1) machine prediction and control learning in a linear walk, and 2) a prediction learning machine interacting with a human participant in a virtual reality environment. Our results showcase the speed of learning for Pavlovian signalling, the impact that different temporal representations do (and do not) have on agent-agent coordination, and how temporal aliasing impacts agent-agent and human-agent interactions differently. As a main contribution, we establish Pavlovian signalling as a natural bridge between fixed signalling paradigms and fully adaptive communication learning. Our results therefore point to an actionable, constructivist path towards continual communication learning between reinforcement learning agents, with potential impact in a range of real-world settings.

The Frost Hollow Experiments: Pavlovian Signalling as a Path to Coordination and Communication Between Agents

Pavlovian Signalling with General Value Functions in Agent-Agent Temporal Decision Making

Learned human-agent decision-making, communication and joint action in a virtual reality environment

Multi-Agent Curricula and Emergent Implicit Signaling

Continually Learned Pavlovian Signalling Without Forgetting for Human-in-the-Loop Robotic Control

Prosody as a Teaching Signal for Agent Learning: Exploratory Studies and Algorithmic Implications

Interpretable Learned Emergent Communication for Human-Agent Teams

Effective Communications: A Joint Learning and Communication Framework for Multi-Agent Reinforcement Learning over Noisy Channels

Verco: Learning Coordinated Verbal Communication for Multi-agent Reinforcement Learning

Improved Cooperation by Exploiting a Common Signal

Learning Effective Communication for Cooperative Pursuit with Multi-Agent Reinforcement Learning

Improving Coordination in Small-Scale Multi-Agent Deep Reinforcement Learning through Memory-driven Communication

Signal Instructed Coordination in Cooperative Multi-agent Reinforcement Learning

Enhancing Multi-Agent Coordination through Common Operating Picture Integration

Mixed Cooperative-Competitive Communication Using Multi-Agent Reinforcement Learning

Learning to communicate about shared procedural abstractions

Assessing Human Interaction in Virtual Reality With Continually Learning Prediction Agents Based on Reinforcement Learning Algorithms: A Pilot Study

R-MADDPG for Partially Observable Environments and Limited Communication

Modeling Communication to Coordinate Perspectives in Cooperation

Learning to Coordinate without Communication under Incomplete Information

The Dynamics of Reinforcement Social Learning in Networked Cooperative Multiagent Systems