Abstract:Over the last ten years, social media has become a crucial data source for businesses and researchers, providing a space where people can express their opinions and emotions. To analyze this data and classify emotions and their polarity in texts, natural language processing (NLP) techniques such as emotion analysis (EA) and sentiment analysis (SA) are employed. However, the effectiveness of these tasks using machine learning (ML) and deep learning (DL) methods depends on large labeled datasets, which are scarce in languages like Spanish. To address this challenge, researchers use data augmentation (DA) techniques to artificially expand small datasets. This study aims to investigate whether DA techniques can improve classification results using ML and DL algorithms for sentiment and emotion analysis of Spanish texts. Various text manipulation techniques were applied, including transformations, paraphrasing (back-translation), and text generation using generative adversarial networks, to small datasets such as song lyrics, social media comments, headlines from national newspapers in Chile, and survey responses from higher education students. The findings show that the Convolutional Neural Network (CNN) classifier achieved the most significant improvement, with an 18% increase using the Generative Adversarial Networks for Sentiment Text (SentiGan) on the Aggressiveness (Seriousness) dataset. Additionally, the same classifier model showed an 11% improvement using the Easy Data Augmentation (EDA) on the Gender-Based Violence dataset. The performance of the Bidirectional Encoder Representations from Transformers (BETO) also improved by 10% on the back-translation augmented version of the October 18 dataset, and by 4% on the EDA augmented version of the Teaching survey dataset. These results suggest that data augmentation techniques enhance performance by transforming text and adapting it to the specific characteristics of the dataset. Through experimentation with various augmentation techniques, this research provides valuable insights into the analysis of subjectivity in Spanish texts and offers guidance for selecting algorithms and techniques based on dataset features.

Detection of violent speech against women in Mexican tweets using an active learning approach

Domain-adaptive pre-training on a BERT model for the automatic detection of misogynistic tweets in Spanish

Machine Unlearning reveals that the Gender-based Violence Victim Condition can be detected from Speech in a Speaker-Agnostic Setting

Semi-Automatic Dataset Annotation Applied to Automatic Violent Message Detection

Machine Learning to study the impact of gender-based violence in the news media

Anti-Sexism Alert System: Identification of Sexist Comments on Social Media Using AI Techniques

Study of violence against women and its characteristics through the application of text mining techniques

Design and development of an assisted communication system, using machine learning and natural language processing, for the detection of fake news on Twitter

A Multichannel Deep Learning Framework for Cyberbullying Detection on Social Media

Breaking the Silence Detecting and Mitigating Gendered Abuse in Hindi, Tamil, and Indian English Online Spaces

A Streaming Machine Learning Framework for Online Aggression Detection on Twitter

Machine learning for risk assessment in gender-based crime

Dominant Set-based Active Learning for Text Classification and its Application to Online Social Media

Deep Learning Approaches for Detecting Adversarial Cyberbullying and Hate Speech in Social Networks

Leveraging external resources for offensive content detection in social media

Automated Detection of Cyberbullying Against Women and Immigrants and Cross-domain Adaptability

A large-scale crowdsourced analysis of abuse against women journalists and politicians on Twitter

Guide for the application of the data augmentation approach on sets of texts in Spanish for sentiment and emotion analysis

A Hybrid CRNN Model for Multi-Class Violence Detection in Text and Video

Improving automatic cyberbullying detection in social network environments by fine-tuning a pre-trained sentence transformer language model

Misogynistic Tweet Detection: Modelling CNN with Small Datasets