Learning with sparse reward in a gap junction network inspired by the insect mushroom body
Tianqi Wei,Qinghai Guo,Barbara Webb
DOI: https://doi.org/10.1371/journal.pcbi.1012086
2024-05-24
PLoS Computational Biology
Abstract:Animals can learn in real-life scenarios where rewards are often only available when a goal is achieved. This 'distal' or 'sparse' reward problem remains a challenge for conventional reinforcement learning algorithms. Here we investigate an algorithm for learning in such scenarios, inspired by the possibility that axo-axonal gap junction connections, observed in neural circuits with parallel fibres such as the insect mushroom body, could form a resistive network. In such a network, an active node represents the task state, connections between nodes represent state transitions and their connection to actions, and current flow to a target state can guide decision making. Building on evidence that gap junction weights are adaptive, we propose that experience of a task can modulate the connections to form a graph encoding the task structure. We demonstrate that the approach can be used for efficient reinforcement learning under sparse rewards, and discuss whether it is plausible as an account of the insect mushroom body. Learning in situations where reward is only rarely encountered is difficult. It is hard to discover the right sequence of actions when most actions, most of the time, provide no apparent progress towards a goal. Inspired by a neural circuit in the insect brain, and using direct electrical connections between neurons as well as synaptic connections, we present a new algorithm for learning. The model represents the states of the world with nodes and an electrical connection between two nodes is strengthened when the two corresponding states occur consecutively. The connections between nodes can also become associated to output actions that correlate with (hence are assumed to cause) transitions between states. When a particular goal is chosen or associated with a reward, for example, the target location in a navigation task, a flow of electrical current through the nodes will find the shortest path from the present state to the goal state and trigger the appropriate actions.
biochemical research methods,mathematical & computational biology