Abstract:Extracting web content is to obtain the required data embedded in web pages, usually including structured records, such as product information, and text content, such as news. Web pages use a large number of HTML tags to organize and to present various information. Both knowing little about the structures of web pages and mixing kinds of information in web pages are making the extraction process very challenging to guarantee extraction performance and extraction adaptability. This study proposes a unified web content extraction framework that can be applied in various web environments to extract both structured records and text content. First, we construct a characteristic container to hold kinds of characteristics related with extraction objectives, including visual text information, content semantics(instead of HTML tag semantics), web page structures, etc. Second, the above characteristics are integrated into an extraction framework for extraction decisions on different web sites. Especially, we put forward different strategies, path aggregation for extracting text content and HMM model for structured records, to locate the extraction area by exploiting both those extraction characteristics. Comparative experiments on multiple web sites with popular extraction methods, including CETR, CETD and CNBE, show that our proposed extraction method can provide better extraction precision and extraction adaptability.

Adaptively Extracting Structured Data from Web Pages

Exploiting Multi-Category Characteristics and Unified Framework to Extract Web Content

Web Information Segmentation Method Based on DOM Structure Tree

A dynamic learning framework to thoroughly extract structured data from web pages without human efforts

Automatic Extraction of Semi-structured Web Data

Adaptive Web Information Extraction Based on DOM Tree

Data extraction from web pages based on structural-semantic entropy.

Dom Semantic Expansion-Based Extraction Of Topical Information From Web Pages

An Algorithm on Web Article Automatic Extraction Based on DOM Structure

Using Structured Tokens To Identify Webpages For Data Extraction

The research and implementation of web information extraction technology based on multi-level pages

A robust approach of automatic web data record extraction

Automatic Web information extraction based on page clustering

Automatic Data Extraction from Data-Rich Web Pages

Web Content Extraction & Its Data Management Method

The Technology of Extracting Content Information from Web Page Based on DOM Tree

Web Entities Extraction Based on Semi-Structured Semantic Database.

Extraction of Relevant Components Using Shallow Structure of HTML Documents.

Dom based extraction of topical information from web pages

A template-based method for theme information extraction from web pages

An Adaptive Web Information Extraction Approach Based on Stu-Dom Tree