Information Extraction in Domain and Generic Documents: Findings from Heuristic-based and Data-driven Approaches

Shiyu Yuan,Carlo Lipizzi

2023-07-01

Abstract:Information extraction (IE) plays very important role in natural language processing (NLP) and is fundamental to many NLP applications that used to extract structured information from unstructured text data. Heuristic-based searching and data-driven learning are two main stream implementation approaches. However, no much attention has been paid to document genre and length influence on IE tasks. To fill the gap, in this study, we investigated the accuracy and generalization abilities of heuristic-based searching and data-driven to perform two IE tasks: named entity recognition (NER) and semantic role labeling (SRL) on domain-specific and generic documents with different length. We posited two hypotheses: first, short documents may yield better accuracy results compared to long documents; second, generic documents may exhibit superior extraction outcomes relative to domain-dependent documents due to training document genre limitations. Our findings reveals that no single method demonstrated overwhelming performance in both tasks. For named entity extraction, data-driven approaches outperformed symbolic methods in terms of accuracy, particularly in short texts. In the case of semantic roles extraction, we observed that heuristic-based searching method and data-driven based model with syntax representation surpassed the performance of pure data-driven approach which only consider semantic information. Additionally, we discovered that different semantic roles exhibited varying accuracy levels with the same method. This study offers valuable insights for downstream text mining tasks, such as NER and SRL, when addressing various document features and genres.

Computation and Language

What problem does this paper attempt to address?

This paper attempts to address the performance and generalization ability of Information Extraction (IE) tasks across different domains (specific and general) and document lengths. Specifically, the study focuses on the performance differences between heuristic-based methods and data-driven methods in two tasks: Named Entity Recognition (NER) and Semantic Role Labeling (SRL). The paper proposes two hypotheses: 1. Short documents may yield better accuracy results than long documents. 2. Due to the limitations of training document types, general documents may exhibit better extraction performance than domain-dependent documents. By comparing the performance of these two methods on different types of documents, the study aims to fill the gap in the existing literature regarding the impact of document type and length on information extraction tasks, thereby providing valuable insights for downstream text mining tasks.

Information Extraction in Domain and Generic Documents: Findings from Heuristic-based and Data-driven Approaches

A Survey of Document-Level Information Extraction

Knowledge Extraction in Low-Resource Scenarios: Survey and Perspective

Assessing the Performance of Chinese Open Source Large Language Models in Information Extraction Tasks

A Study of Recent Contributions on Information Extraction

Data-efficient End-to-end Information Extraction for Statistical Legal Analysis

Research on Information Extraction:A Survey

Large Language Models for Generative Information Extraction: A Survey

Ontologies and Information Extraction

Information Extraction in Illicit Web Domains

Generalisation in Named Entity Recognition: A Quantitative Analysis

Natural language processing for word sense disambiguation and information extraction

Data-Efficient Information Extraction from Form-Like Documents

Dave: Extracting Domain Attributes And Values From Text Corpus

QA4IE: A Question Answering Based System for Document-Level General Information Extraction

Single Document Keyword Extraction for Internet News Articles

Ontology-driven Information Extraction

Open Information Extraction: A Review of Baseline Techniques, Approaches, and Applications

A Survey on Open Information Extraction from Rule-based Model to Large Language Model (meta)

Document-level Entity-based Extraction as Template Generation

Information extraction from electronic medical documents: state of the art and future research directions