Abstract:Multilingual transformers (XLM, mT5) have been shown to have remarkable transfer skills in zero-shot settings. Most transfer studies, however, rely on automatically translated resources (XNLI, XQuAD), making it hard to discern the particular linguistic knowledge that is being transferred, and the role of expert annotated monolingual datasets when developing task-specific models. We investigate the cross-lingual transfer abilities of XLM-R for Chinese and English natural language inference (NLI), with a focus on the recent large-scale Chinese dataset OCNLI. To better understand linguistic transfer, we created 4 categories of challenge and adversarial tasks (totaling 17 new datasets) for Chinese that build on several well-known resources for English (e.g., HANS, NLI stress-tests). We find that cross-lingual models trained on English NLI do transfer well across our Chinese tasks (e.g., in 3/4 of our challenge categories, they perform as well/better than the best monolingual models, even on 3/5 uniquely Chinese linguistic phenomena such as idioms, pro drop). These results, however, come with important caveats: cross-lingual models often perform best when trained on a mixture of English and high-quality monolingual NLI data (OCNLI), and are often hindered by automatically translated resources (XNLI-zh). For many phenomena, all models continue to struggle, highlighting the need for our new diagnostics to help benchmark Chinese and cross-lingual models. All new datasets/code are released at <a class="link-external link-https" href="https://github.com/huhailinguist/ChineseNLIProbing" rel="external noopener nofollow">this https URL</a>.

Cross-lingual Opinion Analysis Via Negative Transfer Detection.

Investigating cross-lingual training for offensive language detection

A Survey on Negative Transfer

A cross-lingual transfer learning method for online COVID-19-related hate speech detection

Cross-lingual Offensive Language Detection: A Systematic Review of Datasets, Transfer Approaches and Challenges

Cross-Lingual Propagation for Deep Sentiment Analysis

Cross-Cultural Transfer Learning for Chinese Offensive Language Detection

Mitigating Negative Transfer with Task Awareness for Sexism, Hate Speech, and Toxic Language Detection

Characterizing and Avoiding Negative Transfer

Cross-lingual Transfer Can Worsen Bias in Sentiment Analysis

Cross-lingual offensive speech identification with transfer learning for low-resource languages

Mitigating Negative Style Transfer in Hybrid Dialogue System

Investigating Transfer Learning in Multilingual Pre-trained Language Models through Chinese Natural Language Inference

Analysing Cross-Lingual Transfer in Low-Resourced African Named Entity Recognition

Transferring Audio Deepfake Detection Capability Across Languages

An Efficient Approach for Studying Cross-Lingual Transfer in Multilingual Language Models

Choosing Transfer Languages for Cross-Lingual Learning

Evaluating and explaining training strategies for zero-shot cross-lingual news sentiment analysis

On Negative Interference in Multilingual Models: Findings and A Meta-Learning Treatment

CLOpinionMiner: Opinion Target Extraction in a Cross-Language Scenario

Learning multilinguistic knowledge for opinion analysis