Abstract:Real-world choice options have many features or attributes, whereas the reward outcome from those options only depends on a few features/attributes. It has been shown that humans learn and combine feature-based with more complex conjunction-based learning to tackle challenges of learning in complex reward environments. However, it is unclear how different learning strategies interact to determine what features should be attended and control choice behavior, and how ensuing attention modulates future learning and/or choice. To address these questions, we examined human behavior during a three-dimensional learning task in which reward outcomes for different stimuli could be predicted based on a combination of an informative feature and conjunction. Using multiple approaches, we first confirmed that choice behavior and reward probabilities estimated by participants were best described by a model that learned the predictive values of both the informative feature and the informative conjunction. In this model, attention was controlled by the difference in these values in a cooperative manner such that attention depended on the integrated feature and conjunction values, and the resulting attention weights modulated learning by increasing the learning rate on attended features and conjunctions. However, there was little effect of attention on decision making. These results suggest that in multidimensional environments, humans direct their attention not only to selectively process reward-predictive attributes, but also to find parsimonious representations of the reward contingencies for more efficient learning.

Contributions of attention to learning in multidimensional reward environments