The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases

Dipto Das,Shion Guha,Jed Brubaker,Bryan Semaan

2024-01-19

Abstract:While colonization has sociohistorically impacted people's identities across various dimensions, those colonial values and biases continue to be perpetuated by sociotechnical systems. One category of sociotechnical systems--sentiment analysis tools--can also perpetuate colonial values and bias, yet less attention has been paid to how such tools may be complicit in perpetuating coloniality, although they are often used to guide various practices (e.g., content moderation). In this paper, we explore potential bias in sentiment analysis tools in the context of Bengali communities that have experienced and continue to experience the impacts of colonialism. Drawing on identity categories most impacted by colonialism amongst local Bengali communities, we focused our analytic attention on gender, religion, and nationality. We conducted an algorithmic audit of all sentiment analysis tools for Bengali, available on the Python package index (PyPI) and GitHub. Despite similar semantic content and structure, our analyses showed that in addition to inconsistencies in output from different tools, Bengali sentiment analysis tools exhibit bias between different identity categories and respond differently to different ways of identity expression. Connecting our findings with colonially shaped sociocultural structures of Bengali communities, we discuss the implications of downstream bias of sentiment analysis tools.

Computation and Language,Computers and Society,Human-Computer Interaction,Machine Learning

What problem does this paper attempt to address?

The problem this paper attempts to address is whether sentiment analysis tools in Natural Language Processing (NLP) perpetuate colonial influences when dealing with the Bengali language, and whether these tools exhibit biases when evaluating different identity categories such as gender, religion, and nationality. Specifically, the authors focus on the following aspects: 1. **What are the differences in sentiment scores for specific identities across different tools** (RQ1.a)? 2. **What are the differences in sentiment scores for explicit and implicit identity expressions** (RQ1.b)? 3. **Do Bengali Sentiment Analysis (BSA) tools exhibit biases in the categories of gender, religion, and nationality** (RQ2.a)? 4. **What is the relationship between the biases of the tools and the background of the developers** (RQ2.b)? Through these questions, the authors aim to explore and reveal whether and how these tools reactivate social biases shaped by colonialism across different identity dimensions. This not only helps in understanding social and technical biases in technological systems but also provides important references and suggestions for future research.

The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases

A Material Lens on Coloniality in NLP

Exploring Bengali Religious Dialect Biases in Large Language Models with Evaluation Perspectives

Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious Bias

Global Voices, Local Biases: Socio-Cultural Prejudices across Languages

Gender Bias Mitigation for Bangla Classification Tasks

An Empirical Study of Gendered Stereotypes in Emotional Attributes for Bangla in Multilingual Large Language Models

Exploring the perception of ethnic migrants in Kolkata, India: A comparative study using sentiment analysis

Preparing Bengali-English Code-Mixed Corpus for Sentiment Analysis of Indian Languages

Milestones in Bengali Sentiment Analysis leveraging Transformer-models: Fundamentals, Challenges and Future Directions

An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla

Mapping the Multilingual Margins: Intersectional Biases of Sentiment Analysis Systems in English, Spanish, and Arabic

Sentiment Polarity Detection on Bengali Book Reviews Using Multinomial Naive Bayes

Enhancing Sentiment Analysis in Bengali Texts: A Hybrid Approach Using Lexicon-Based Algorithm and Pretrained Language Model Bangla-BERT

Analyzing Roles of Classifiers and Code-Mixed factors for Sentiment Identification

Analysis and Mitigation of Religion Bias in Indonesian Natural Language Processing Datasets

Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models

Mitigating Gender Stereotypes in Hindi and Marathi

Reimagining Communities through Transnational Bengali Decolonial Discourse with YouTube Content Creators

Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting

Sentiment analysis in Bengali via transfer learning using multi-lingual BERT