Multi-Layer Features Based Personalized Spam Filtering

Weiran Xu,Zhanyi Wang,Dongxin Liu,Jun Guo,Rile Hu
DOI: https://doi.org/10.1109/icnidc.2009.5360803
2009-01-01
Abstract:In this paper, we face a new challenge that the filter is expected to converge much faster, e.g. within 10 labeled SMSs or less. Topic model based dimension reduction can minimize the structural risk with limited training data. But dimension reduction will go against the completeness of feature space. It is very difficult to obtain the convergence rate and the completeness at the same time only by one kind of feature. This paper uses supervised dual-PLSA for Dimensionality Reduction and presents a multi-layer features model, which employs two layer features and adopts a novel method to combine them. Experiments show that multi-layer features model have the best performance.
What problem does this paper attempt to address?