A General Multi-relational Classification Approach Using Feature Generation and Selection.

Miao Zou,Tengjiao Wang,Hongyan Li,Dongqing Yang
DOI: https://doi.org/10.1007/978-3-642-17313-4_3
2010-01-01
Abstract:Multi-relational classification is an important data mining task, since much real world data is organized in multiple relations. The major challenges come from, firstly, the large high dimensional search spaces due to many attributes in multiple relations and, secondly, the high computational cost in feature selection and classifier construction due to the high complexity in the structure of multiple relations. The existing approaches mainly use the inductive logic programming (ILP) techniques to derive hypotheses or extract features for classification. However, those methods often are slow and sometimes cannot provide enough information to build effective classifiers. In this paper, we develop a general approach for accurate and fast multi-relational classification using feature generation and selection. Moreover, we propose a novel similarity-based feature selection method for multi-relational classification. An extensive performance study on several benchmark data sets indicates that our approach is accurate, fast and highly scalable.
What problem does this paper attempt to address?