Abstract:Multiple entities in a document generally exhibit complex inter-sentence relations, and cannot be well handled by existing relation extraction (RE) methods that typically focus on extracting intxa-sentence relations for single entity pairs. In order to accelerate the research on document-level RE, we introduce DocRED, a new datasct constructed from Wikipedia and Wikidata with three features: (1) DocRED annotates both named entities and relations, and is the largest human-annotated dataset for document-level RE from plain text; (2) DocRED requires reading multiple sentences in a document to extract entities and infer their relations by synthesizing all information of the document; (3) along with the human-annotated data, we also offer large-scale distantly supervised data, which enables DocRED to be adopted for both supervised and weakly supervised scenarios. In order to verify the challenges of document-level RE, we implement recent state-of-the-art methods for RE and conduct a thorough evaluation of these methods on DocRED. Empirical results show that DocRED is challenging for existing RE methods, which indicates that document-level RE remains an open problem and requires further efforts. Based on the detailed analysis on the experiments, we discuss multiple promising directions for future research. We make DocRED and the code for our baselines publicly available at https://github.com/thun1p/DocRED.

Manual Evaluation Matters: Reviewing Test Protocols of Distantly Supervised Relation Extraction.

Rescue Implicit and Long-tail Cases: Nearest Neighbor Relation Extraction

Towards Accurate and Consistent Evaluation: A Dataset for Distantly-Supervised Relation Extraction

Revisiting Distant Supervision for Relation Extraction

Towards Realistic Low-resource Relation Extraction: A Benchmark with Empirical Baseline Study

Did the Models Understand Documents? Benchmarking Models for Language Understanding in Document-Level Relation Extraction

Revisiting the Negative Data of Distantly Supervised Relation Extraction

On Robustness and Bias Analysis of BERT-Based Relation Extraction

Fine-tuning Pre-Trained Transformer Language Models to Distantly Supervised Relation Extraction

GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models

Can We Have Both Fish and Bear's Paw? Improving Performance, Reliability, and both of them for Relation Extraction under Label Shift

Distantly Supervised Relation Extraction Based on Residual Attention and Self Learning

Improving Relation Extraction with Relational Paraphrase Sentences.

A Review on Semi-Supervised Relation Extraction

Can Fine-tuning Pre-trained Models Lead to Perfect NLP? A Study of the Generalizability of Relation Extraction.

Can We Have Both Fish and Bear's Paw?

Docred: A Large-Scale Document-Level Relation Extraction Dataset

What do You Mean by Relation Extraction? A Survey on Datasets and Study on Scientific Relation Classification

Knowledge-Aware And Retrieval-Based Models For Distantly Supervised Relation Extraction

Finding Influential Instances for Distantly Supervised Relation Extraction

LTACL: long-tail awareness contrastive learning for distantly supervised relation extraction