Abstract:Category learning performance is influenced by both the nature of the category's structure and the way category features are processed during learning. Shepard (1964, 1987) showed that stimuli can have structures with features that are statistically uncorrelated (separable) or statistically correlated (integral) within categories. Humans find it much easier to learn categories having separable features, especially when attention to only a subset of relevant features is required, and harder to learn categories having integral features, which require consideration of all of the available features and integration of all the relevant category features satisfying the category rule (Garner, 1974). In contrast to humans, a single hidden layer backpropagation (BP) neural network has been shown to learn both separable and integral categories equally easily, independent of the category rule (Kruschke, 1993). This "failure" to replicate human category performance appeared to be strong evidence that connectionist networks were incapable of modeling human attentional bias. We tested the presumed limitations of attentional bias in networks in two ways: (1) by having networks learn categories with exemplars that have high feature complexity in contrast to the low dimensional stimuli previously used, and (2) by investigating whether a Deep Learning (DL) network, which has demonstrated humanlike performance in many different kinds of tasks (language translation, autonomous driving, etc.), would display human-like attentional bias during category learning. We were able to show a number of interesting results. First, we replicated the failure of BP to differentially process integral and separable category structures when low dimensional stimuli are used (Garner, 1974; Kruschke, 1993). Second, we show that using the same low dimensional stimuli, Deep Learning (DL), unlike BP but similar to humans, learns separable category structures more quickly than integral category structures. Third, we show that even BP can exhibit human like learning differences between integral and separable category structures when high dimensional stimuli (face exemplars) are used. We conclude, after visualizing the hidden unit representations, that DL appears to extend initial learning due to feature development thereby reducing destructive feature competition by incrementally refining feature detectors throughout later layers until a tipping point (in terms of error) is reached resulting in rapid asymptotic learning.

A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark

Re-examining Lexical and Semantic Attention: Dual-view Graph Convolutions Enhanced BERT for Academic Paper Rating.

Fine-tune BERT with Sparse Self-Attention Mechanism.

SesameBERT: Attention for Anywhere

What Does BERT Look At? An Analysis of BERT's Attention

A Closer Look at How Fine-tuning Changes BERT

Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models

Context-Guided BERT for Targeted Aspect-Based Sentiment Analysis

Syntax Matters: Towards Spoken Language Understanding Via Syntax-Aware Attention

Paying More Attention to Self-attention: Improving Pre-trained Language Models via Attention Guiding

Attention Flows: Analyzing and Comparing Attention Mechanisms in Language Models

Can ChatGPT Understand Too? A Comparative Study on ChatGPT and Fine-tuned BERT

Attentional Bias in Human Category Learning: The Case of Deep Learning

A Study of Cross-Lingual Ability and Language-specific Information in Multilingual BERT

Investigating Novel Verb Learning in BERT: Selectional Preference Classes and Alternation-Based Syntactic Generalization

Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration

Improving BERT-Based Text Classification With Auxiliary Sentence and Domain Knowledge

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

An Exploratory Study on Code Attention in BERT