Abstract:Warning: This paper contains examples of the language that some people may find offensive. Detecting and reducing hateful, abusive, offensive comments is a critical and challenging task on social media. Moreover, few studies aim to mitigate the intensity of hate speech. While studies have shown that context-level semantics are crucial for detecting hateful comments, most of this research focuses on English due to the ample datasets available. In contrast, low-resource languages, like Indian languages, remain under-researched because of limited datasets. Contrary to hate speech detection, hate intensity reduction remains unexplored in high-resource and low-resource languages. In this paper, we propose a novel end-to-end model, HCDIR, for Hate Context Detection, and Hate Intensity Reduction in social media posts. First, we fine-tuned several pre-trained language models to detect hateful comments to ascertain the best-performing hateful comments detection model. Then, we identified the contextual hateful words. Identification of such hateful words is justified through the state-of-the-art explainable learning model, i.e., Integrated Gradient (IG). Lastly, the Masked Language Modeling (MLM) model has been employed to capture domain-specific nuances to reduce hate intensity. We masked the 50\% hateful words of the comments identified as hateful and predicted the alternative words for these masked terms to generate convincing sentences. An optimal replacement for the original hate comments from the feasible sentences is preferred. Extensive experiments have been conducted on several recent datasets using automatic metric-based evaluation (BERTScore) and thorough human evaluation. To enhance the faithfulness in human evaluation, we arranged a group of three human annotators with varied expertise.

Ruddit: Norms of Offensiveness for English Reddit Comments

Measuring Offensive Speech in Online Political Discourse

Norm violation in online communities -- A study of Stack Overflow comments

D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation

Like trainer, like bot? Inheritance of bias in algorithmic content moderation

"It's Not Just Hate'': A Multi-Dimensional Perspective on Detecting Harmful Speech Online

The lack of theory is painful: Modeling Harshness in Peer Review Comments

Trawling for Trolling: A Dataset

OffLanDat: A Community Based Implicit Offensive Language Dataset Generated by Large Language Model Through Prompt Engineering

CoRAL: a Context-aware Croatian Abusive Language Dataset

Mitigating Biases to Embrace Diversity: A Comprehensive Annotation Benchmark for Toxic Language

Offensive Language Detection: A Comparative Analysis

HCDIR: End-to-end Hate Context Detection, and Intensity Reduction model for online comments

Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive

''Fifty Shades of Bias'': Normative Ratings of Gender Bias in GPT Generated English Text

OLID-BR: offensive language identification dataset for Brazilian Portuguese

Measuring the Prevalence of Anti-Social Behavior in Online Communities

DravidianCodeMix: sentiment analysis and offensive language identification dataset for Dravidian languages in code-mixed text

Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Korean Online Hate Speech Dataset for Multilabel Classification: How Can Social Science Improve Dataset on Hate Speech?

Exploratory Data Analysis on Code-mixed Misogynistic Comments