Abstract:Recent research in both academia and industry has validated the effectiveness of provenance graph-based detection for advanced cyber attack detection and investigation. However, analyzing large-scale provenance graphs often results in substantial overhead. To improve performance, existing detection systems implement various optimization strategies. Yet, as several recent studies suggest, these strategies could lose necessary context information and be vulnerable to evasions. Designing a detection system that is efficient and robust against adversarial attacks is an open problem. We introduce Marlin, which approaches cyber attack detection through real-time provenance graph <a class="link-external link-http" href="http://alignment.By" rel="external noopener nofollow">this http URL</a> leveraging query graphs embedded with attack knowledge, Marlin can efficiently identify entities and events within provenance graphs, embedding targeted analysis and significantly narrowing the search space. Moreover, we incorporate our graph alignment algorithm into a tag propagation-based schema to eliminate the need for storing and reprocessing raw logs. This design significantly reduces in-memory storage requirements and minimizes data processing overhead. As a result, it enables real-time graph alignment while preserving essential context information, thereby enhancing the robustness of cyber attack detection. Moreover, Marlin allows analysts to customize attack query graphs flexibly to detect extended attacks and provide interpretable detection results. We conduct experimental evaluations on two large-scale public datasets containing 257.42 GB of logs and 12 query graphs of varying sizes, covering multiple attack techniques and scenarios. The results show that Marlin can process 137K events per second while accurately identifying 120 subgraphs with 31 confirmed attacks, along with only 1 false positive, demonstrating its efficiency and accuracy in handling massive data.

GHunter: A Fast Subgraph Matching Method for Threat Hunting.

ASA: Adversary Situation Awareness Via Heterogeneous Graph Convolutional Networks.

MEGR-APT: A Memory-Efficient APT Hunting System Based on Attack Representation Learning

LogKernel A Threat Hunting Approach Based on Behaviour Provenance Graph and Graph Kernel Clustering

APT-KGL: an Intelligent APT Detection System Based on Threat Knowledge and Heterogeneous Provenance Graph Learning

A Hierarchical Approach for Advanced Persistent Threat Detection with Attention-Based Graph Neural Networks

A System for Efficiently Hunting for Cyber Threats in Computer Systems Using Threat Intelligence

Enabling Efficient Cyber Threat Hunting with Cyber Threat Intelligence

threaTrace: Detecting and Tracing Host-based Threats in Node Level Through Provenance Graph Learning

TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph

A Heterogeneous Graph Learning Model for Cyber-Attack Detection

Combating Advanced Persistent Threats: Challenges and Solutions

TREC: APT Tactic / Technique Recognition via Few-Shot Provenance Subgraph Learning

Sequence Feature Extraction-Based APT Attack Detection Method with Provenance Graphs

APTKG: Constructing Threat Intelligence Knowledge Graph from Open-Source APT Reports Based on Deep Learning

FPGAA: A Multi-Feature Provenance Graph for the Accurate Alert System

Detecting APT-Exploited Processes through Semantic Fusion and Interaction Prediction

Detecting Unknown Threat Based on Continuous-Time Dynamic Heterogeneous Graph Network

Marlin: Knowledge-Driven Analysis of Provenance Graphs for Efficient and Robust Detection of Cyber Attacks

Detect Advanced Persistent Threat In Graph-Level Using Competitive AutoEncoder.

Prov2vec: Learning Provenance Graph Representation for Unsupervised APT Detection