Abstract:There has been a steady need to precisely extract structured knowledge from the web (i.e. HTML documents). Given a web page, extracting a structured object along with various attributes of interest (e.g. price, publisher, author, and genre for a book) can facilitate a variety of downstream applications such as large-scale knowledge base construction, e-commerce product search, and personalized recommendation. Considering each web page is rendered from an HTML DOM tree, existing approaches formulate the problem as a DOM tree node tagging task. However, they either rely on computationally expensive visual feature engineering or are incapable of modeling the relationship among the tree nodes. In this paper, we propose a novel transferable method, Simplified DOM Trees for Attribute Extraction (SimpDOM), to tackle the problem by efficiently retrieving useful context for each node by leveraging the tree structure. We study two challenging experimental settings: (i) intra-vertical few-shot extraction, and (ii) cross-vertical fewshot extraction with out-of-domain knowledge, to evaluate our approach. Extensive experiments on the SWDE public dataset show that SimpDOM outperforms the state-of-the-art (SOTA) method by 1.44% on the F1 score. We also find that utilizing knowledge from a different vertical (cross-vertical extraction) is surprisingly useful and helps beat the SOTA by a further 1.37%.

Adaptive Web Information Extraction Based on DOM Tree

Web Information Segmentation Method Based on DOM Structure Tree

An Adaptive Web Information Extraction Approach Based on Stu-Dom Tree

DOM-Based Automatic Extraction of Topical Information from Web Pages

The Technology of Extracting Content Information from Web Page Based on DOM Tree

Attributes extraction of Deep Web query interface based on DOM

Tag Tree Template for Web Information and Schema Extraction.

An Approach of Information Extraction Based on Dom Tree and Weight Value

Web News Pages Extraction Method Based on DOM and Decision Tree

Topic information extraction from Web pages based on tree comparison

Web Information Extraction Based on Repeated Pattern

Web data extraction based on a simplified Dom Tree

An Algorithm on Web Article Automatic Extraction Based on DOM Structure

Simplified DOM Trees for Transferable Attribute Extraction from the Web

A Semantic DOM Approach for Webpage Information Extraction

Application and Design of Web Information Extraction System Based on Pattern Discovery

STUDY AND IMPLEMENTATION OF DYNAMIC WEB INFORMATION EXTRACTION BASED ON TREE MODEL ALGORITHM

Web Pages Information Retrieval Based on Keywords Cluster and Node Instance

Auto-extraction Methods of Web Pagelet

Content Extraction of Web Pages Based on Characteristic Symbols

Template-based Information Automatic Extraction of Web