Abstract:Extracting structured information from unstructured text is crucial for knowledge management and utilization, which is the goal of document-level relation extraction. Existing graph-based methods face issues with information confusion and integration, limiting the reasoning capabilities of the model. To tackle this problem, a dual-stream dynamic graph structural network is proposed to model documents from various perspectives. Leveraging the richness of document information, a static document heterogeneous graph is constructed. A dynamic heterogeneous document graph is then induced based on this foundation to facilitate global information aggregation for entity representation learning. Additionally, the static document graph is decomposed into multi-level static semantic graphs, and multi-layer dynamic semantic graphs are further induced, explicitly segregating information from different levels. Information from different streams is effectively integrated via an information integrator. To mitigate the interference of noise during the reasoning process, a noise regularization mechanism is also designed. The experimental results on three extensively utilized publicly accessible datasets for document-level relation extraction demonstrate that our model achieves F1 scores of 62.56%, 71.1%, and 86.9% on the DocRED, CDR, and GDA datasets, respectively, significantly outperforming the baselines. Further analysis also demonstrates the effectiveness of the model in multi-entity scenarios.

Logical Structure Based Semantic Relationship Extraction from Semi-Structured Documents

A Rule-Based Information Extraction System for Human-Readable Semi-Structured Scientific Documents

Research and Implementation of Structure Extraction of Semi-structured Document

Optimized Hierarchy Clustering Based Extraction for Logical Document Structures

Searching Semantic Web Documents Based on RDF Sentences

Semantic Text Mining with Linked Data

Hierarchical Logical Structure Extraction of Book Documents by Analyzing Tables of Contents

Extracting semantic information structures from free text law enforcement data

Dual-stream dynamic graph structure network for document-level relation extraction

List data extraction in semi-structured document

Visualizing Multi-Document Semantics Via Open Domain Information Extraction

Extracting relation information from text documents by exploring various types of knowledge

Semantic Relationship-Oriented Biological Text Retrieval

Reasoning with Latent Structure Refinement for Document-Level Relation Extraction

Learning a Probabilistic Semantic Model from Heterogeneous Social Networks for Relationship Identification

An Ontology-Based Approach to Extracting Semantic Relations from Descriptive Text.

DWSA: An Intelligent Document Structural Analysis Model for Information Extraction and Data Mining

Data extraction from web pages based on structural-semantic entropy.

Structure extraction from PDF-based book documents.

Understanding the Logical and Semantic Structure of Large Documents

Latent table discovery by semantic relationship extraction between unrelated sets of entity sets of structured data sources