Abstract:Human ability to understand language is general, flexible, and robust. In contrast, most NLU models above the word level are designed for a specific task and struggle with out-of-domain data. If we aspire to develop models with understanding beyond the detection of superficial correspondences between inputs and outputs, then it is critical to develop a unified model that can execute a range of linguistic tasks across different domains. To facilitate research in this direction, we present the General Language Understanding Evaluation (GLUE, gluebenchmark.com): a benchmark of nine diverse NLU tasks, an auxiliary dataset for probing models for understanding of specific linguistic phenomena, and an online platform for evaluating and comparing models. For some benchmark tasks, training data is plentiful, but for others it is limited or does not match the genre of the test set. GLUE thus favors models that can represent linguistic knowledge in a way that facilitates sample-efficient learning and effective knowledge-transfer across tasks. While none of the datasets in GLUE were created from scratch for the benchmark, four of them feature privately-held test data, which is used to ensure that the benchmark is used fairly. We evaluate baselines that use ELMo (Peters et al., 2018), a powerful transfer learning technique, as well as state-of-the-art sentence representation models. The best models still achieve fairly low absolute scores. Analysis with our diagnostic dataset yields similarly weak performance over all phenomena tested, with some exceptions.

Benchmarking Meaning Representations in Neural Semantic Parsing

PMB5: Gaining More Insight into Neural Semantic Parsing with Challenging Benchmarks

Neural Semantic Parsing with Extremely Rich Symbolic Meaning Representations

Exploring the Secrets Behind the Learning Difficulty of Meaning Representations for Semantic Parsing.

A Fine-grained Interpretability Evaluation Benchmark for Neural NLP

Evaluating Scoped Meaning Representations

BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data

Align-smatch: A Novel Evaluation Method for Chinese Abstract Meaning Representation Parsing based on Alignment of Concept and Relation

BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing

Revisiting a Pain in the Neck: Semantic Phrase Processing Benchmark for Language Models

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Discourse Representation Structure Parsing for Chinese

A Sketch-Based System for Semantic Parsing

Towards Benchmarking Situational Awareness of Large Language Models:Comprehensive Benchmark, Evaluation and Analysis

Survey on Abstract Meaning Representation

An Interpretability Evaluation Benchmark for Pre-trained Language Models

MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis

A Neural Transition-Based Approach for Semantic Dependency Graph Parsing

SciRepEval: A Multi-Format Benchmark for Scientific Document Representations

Benchmarking Foundation Models with Language-Model-as-an-Examiner