Investigating Feature and Model Importance in Android Malware Detection: An Implemented Survey and Experimental Comparison of ML-Based Methods

Ali Muzaffar,Hani Ragab Hassen,Hind Zantout,Michael A Lones

2024-08-26

Abstract:The popularity of Android means it is a common target for malware. Over the years, various studies have found that machine learning models can effectively discriminate malware from benign applications. However, as the operating system evolves, so does malware, bringing into question the findings of these previous studies, many of which report very high accuracies using small, outdated, and often imbalanced datasets. In this paper, we reimplement 18 representative past works and reevaluate them using a balanced, relevant, and up-to-date dataset comprising 124,000 applications. We also carry out new experiments designed to fill holes in existing knowledge, and use our findings to identify the most effective features and models to use for Android malware detection within a contemporary environment. We show that high detection accuracies (up to 96.8%) can be achieved using features extracted through static analysis alone, yielding a modest benefit (1%) from using far more expensive dynamic analysis. API calls and opcodes are the most productive static and TCP network traffic provide the most predictive dynamic features. Random forests are generally the most effective model, outperforming more complex deep learning approaches. Whilst directly combining static and dynamic features is generally ineffective, ensembling models separately leads to performances comparable to the best models but using less brittle features.

Machine Learning,Cryptography and Security

What problem does this paper attempt to address?

The paper aims to address the following issues: 1. **Importance Analysis of Features and Models**: Reimplement and evaluate past research work to determine which static and dynamic features and machine learning models are most effective for Android malware detection in the current Android ecosystem. The study found that features extracted through static analysis (such as API calls and opcodes) can achieve a high detection accuracy of up to 96.8%, and the Random Forest model is generally more effective than complex deep learning methods. 2. **Dataset Update and Expansion**: Construct a balanced, relevant, and up-to-date dataset containing 124,000 application samples (62,000 benign apps and 62,000 malicious apps) to evaluate the effectiveness of various Android anti-malware methods. This helps to overcome the issues of outdated, small-scale, or imbalanced datasets used in previous studies. 3. **Feature Selection and Model Integration**: Fill existing knowledge gaps through experiments and identify the best feature selection algorithms and model combination methods. An integrated method combining optimal static and dynamic models is proposed, achieving an accuracy of 97.8% on contemporary datasets. In summary, the paper focuses on systematically comparing and analyzing the most effective feature and model selection strategies in the modern Android operating system environment to improve the accuracy and robustness of malware detection.

Investigating Feature and Model Importance in Android Malware Detection: An Implemented Survey and Experimental Comparison of ML-Based Methods

Revisiting Static Feature-Based Android Malware Detection

Effective and Explainable Detection of Android Malware Based on Machine Learning Algorithms

An Android Malware Detection System Based on Machine Learning

A Hybrid Model for Android Malware Detection

A Systematic Overview of Android Malware Detection

Research of Malware Detection Approach for Android

Exploring Feature Extraction and ELM in Malware Detection for Android Devices.

An End-to-end Model for Android Malware Detection

Novel Android Malware Detection Method Based on Multi-dimensional Hybrid Features Extraction and Analysis

Two Effective Methods to Detect Mobile Malware

An Android Malware Detection Method Using Multi-Feature and MobileNet.

Dynamic detection of mobile malware using smartphone data and machine learning

Malware Detection Techniques by Mining Massive Behavioral Data of Mobile Apps

A Novel Approach for Mobile Malware Classification and Detection in Android Systems.

An Effective Deep Learning Scheme for Android Malware Detection Leveraging Performance Metrics and Computational Resources

A Hybrid Analysis-Based Approach to Android Malware Family Classification

An Efficient Android Malware Detection System Based on Method-Level Behavioral Semantic Analysis.

An Adaptive Semi-Supervised Deep Learning-Based Framework for the Detection of Android Malware.

Artificial Intelligence Algorithms for Malware Detection in Android-Operated Mobile Devices

DroidDet: Effective and Robust Detection of Android Malware Using Static Analysis along with Rotation Forest Model.