Abstract:Outlier detection plays an important role in the pre-treatment of sequential datasets to obtain pure valuable data. This paper proposes an outlier detection scheme for dynamical sequential datasets. First, the conception of forward outlier factor(FOF) and backward outlier factor(BOF) are employed to measure an object's similarity shared with its sequentially adjacent objects. The object that shows no similarity with its sequential neighbors is labeled as suspicious outliers, which will be treated subsequently to judge whether it is really an outlier in the dataset. Second, the sequentially adjacent suspicious outliers are defined as suspicious outlier series(SOS), then the expected path representing the ideal transition path through the suspicious outliers in the SOS and the measured path representing the real path through all the objects in the SOS are employed, and the ratio of the length of the expected path to that of the measured path indicates whether there exist outliers in the SOS. Third, in the case that there exist outliers in the SOS, if there are N suspicious outliers in the SOS, then 2(N) - 2 remaining path will be generated by removing k(0 < k < N) suspicious outliers and sequentially connecting the remaining ones. The dynamical sequential outlier factor(DSOF) is employed to represent the ratio of the length of measured path of the considered remaining path to the that of the the expected path of the corresponding SOS, and the degree of the objects removed in a remaining path being outliers is indicated by the DSOF. The proposed outlier detection scheme is conducted from a dynamical perspective, and breaks the tight relation between being an outlier and being not similar with adjacent objects. Experiments are conducted to evaluate the effectiveness of the proposed scheme, and the experimental results verify that the proposed scheme has higher detection quality for sequential dataset. In addition, the proposed outlier detection scheme is not dependent on the size of dataset and needs no prior information about the distribution of the data.

Outlier Detection as Instance Selection Method for Feature Selection in Time Series Classification

Sample Weighting: an Inherent Approach for Outlier Suppressing Discriminant Analysis

Fairness-aware Outlier Ensemble

HIERVAR: A Hierarchical Feature Selection Method for Time Series Analysis

An iterative approach to unsupervised outlier detection using ensemble method and distance-based data filtering

A method for outlier detection based on cluster analysis and visual expert criteria

Outlier Ranking in Large-Scale Public Health Streams

Simultaneous feature selection and outlier detection with optimality guarantees

Feature Selection Based on Intrusive Outliers Rather Than All Instances

Sparse Modeling-Based Sequential Ensemble Learning for Effective Outlier Detection in High-Dimensional Numeric Data.

Analysis and comparison of feature selection methods towards performance and stability

Discriminative feature selection with directional outliers correcting for data classification

Outliers Learning And Its Applications

Outlier Detection and Spatial Analysis Algorithms

Outlier Detection Method for Time Series Based on the Rate of Signal Change

An Evaluation of Classification and Outlier Detection Algorithms

An Outlier Detection Scheme for Dynamical Sequential Datasets.

Sparsity-based Feature Selection for Anomalous Subgroup Discovery

Relevancy contemplation in medical data analytics and ranking of feature selection algorithms

Outlier detection using conditional information entropy and rough set theory

Feature Selection: A Data Perspective