Abstract:Previous work has established that a person's demographics and speech style affect how well speech processing models perform for them. But where does this bias come from? In this work, we present the Speech Embedding Association Test (SpEAT), a method for detecting bias in one type of model used for many speech tasks: pre-trained models. The SpEAT is inspired by word embedding association tests in natural language processing, which quantify intrinsic bias in a model's representations of different concepts, such as race or valence (something's pleasantness or unpleasantness) and capture the extent to which a model trained on large-scale socio-cultural data has learned human-like biases. Using the SpEAT, we test for six types of bias in 16 English speech models (including 4 models also trained on multilingual data), which come from the wav2vec 2.0, HuBERT, WavLM, and Whisper model families. We find that 14 or more models reveal positive valence (pleasantness) associations with abled people over disabled people, with European-Americans over African-Americans, with females over males, with U.S. accented speakers over non-U.S. accented speakers, and with younger people over older people. Beyond establishing that pre-trained speech models contain these biases, we also show that they can have real world effects. We compare biases found in pre-trained models to biases in downstream models adapted to the task of Speech Emotion Recognition (SER) and find that in 66 of the 96 tests performed (69%), the group that is more associated with positive valence as indicated by the SpEAT also tends to be predicted as speaking with higher valence by the downstream model. Our work provides evidence that, like text and image-based models, pre-trained speech based-models frequently learn human-like biases. Our work also shows that bias found in pre-trained models can propagate to the downstream task of SER.

Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?

Computational Modeling of Stereotype Content in Text

Stereotype and Skew: Quantifying Gender Bias in Pre-trained and Fine-tuned Language Models

HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection

Are Models Biased on Text without Gender-related Language?

Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling

Sociolectal Analysis of Pretrained Language Models

Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution

StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language Models

Quantifying Stereotypes in Language

Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach

Hate Speech Classifiers Learn Human-Like Social Stereotypes

Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale

Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion Recognition

Towards Auditing Large Language Models: Improving Text-based Stereotype Detection

Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models

A Taxonomy of Stereotype Content in Large Language Models

Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content Model

Exposing Bias in Online Communities through Large-Scale Language Models

From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Dialect prejudice predicts AI decisions about people's character, employability, and criminality