Abstract:System logs can record the system status and important events during system operation in detail. Detecting anomalies in the system logs is a common method for modern large-scale distributed systems. Yet threshold-based classification models used for anomaly detection output only two values: normal or abnormal, which lacks probability of estimating whether the prediction results are correct. In this paper, a statistical learning algorithm Venn-Abers predictor is adopted to evaluate the confidence of prediction results in the field of system log anomaly detection. It is able to calculate the probability distribution of labels for a set of samples and provide a quality assessment of predictive labels to some extent. Two Venn-Abers predictors LR-VA and SVM-VA have been implemented based on Logistic Regression and Support Vector Machine, respectively. Then, the differences among different algorithms are considered so as to build a multimodel fusion algorithm by Stacking. And then a Venn-Abers predictor based on the Stacking algorithm called Stacking-VA is implemented. The performances of four types of algorithms (unimodel, Venn-Abers predictor based on unimodel, multimodel, and Venn-Abers predictor based on multimodel) are compared in terms of validity and accuracy. Experiments are carried out on a log dataset of the Hadoop Distributed File System (HDFS). For the comparative experiments on unimodels, the results show that the validities of LR-VA and SVM-VA are better than those of the two corresponding underlying models. Compared with the underlying model, the accuracy of the SVM-VA predictor is better than that of LR-VA predictor, and more significantly, the recall rate increases from 81% to 94%. In the case of experiments on multiple models, the algorithm based on Stacking multimodel fusion is significantly superior to the underlying classifier. The average accuracy of Stacking-VA is larger than 0.95, which is more stable than the prediction results of LR-VA and SVM-VA. Experimental results show that the Venn-Abers predictor is a flexible tool that can make accurate and valid probability predictions in the field of system log anomaly detection.

Improving Log Anomaly Detection Via Spatial Pooling: Combining Spclassifier with Ensemble Method

Improving Log Anomaly Detection Via Spatial Pooling: Combining SPClassifier with Ensemble Method

An Anomaly Detection Approach of Part-of-Speech Log Sequence Via Population Based Training

Detection of Anomalies in Multivariate Time Series Using Ensemble Techniques

Log anomaly detection method based on CNN and LSTM fusion

Research on An Ensemble Anomaly Detection Algorithm

Experience Report: Deep Learning-based System Log Analysis for Anomaly Detection

Improving Performance of Log Anomaly Detection with Semantic and Time Features Based on BiLSTM-Attention

Try with Simpler -- An Evaluation of Improved Principal Component Analysis in Log-based Anomaly Detection

Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework

Anomaly Detection Using XGBoost Ensemble of Deep Neural Network Models

Improving imbalance classification via ensemble learning based on two-stage learning

New hybrid ensemble method for anomaly detection in data science

OneLog: Towards End-to-End Training in Software Log Anomaly Detection

Enhancing Anomaly Detection in Pedestrian Walkways using Improved Sparrow Search Algorithm with Parallel Features Fusion Model

Valid Probabilistic Anomaly Detection Models for System Logs

Distributed system anomaly detection using deep learning‐based log analysis

AIOps-Driven Enhancement of Log Anomaly Detection in Unsupervised Scenarios

LogCAE: an Approach for Log-based Anomaly Detection with Active Learning and Contrastive Learning

An Improved SVM Method in Anomaly Detection

Robust and accurate performance anomaly detection and prediction for cloud applications: a novel ensemble learning-based framework