PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 29, 2026Computers, materials & continua/Computers, materials & continua (Print)0 citationsOpen Access

A Hybrid Self-Supervised Learning Framework for Advanced Persistent Threat Detection

View Full Paper
MAMarwan Ali Albahar

Key Points

  • The aim is to enhance anomaly detection for advanced persistent threats using a self-supervised learning approach.
  • Developed a framework called Self-Training Adaptive Graph Encoder (stage)
  • Utilized Graph Convolutional Networks for embedding learning
  • Incorporated a memory augmented attention module to capture benign patterns
  • Combined contrastive learning with one-class Support Vector Data Description
  • Fused neural embeddings with classic detection methods for inference.
  • Achieved 95% recall with 0% false positive rate on the StreamSpot dataset
  • Attained AUC of 0.998 on both StreamSpot and Wget datasets
  • Maintained 100% recall and 96% precision with a 4% false positive rate on the Wget dataset
  • Demonstrated effective empirical separability for benign-only detection.

Abstract

Advanced Persistent Threats (APTs) are stealthy cyberattacks that can evade detection in system-level audit logs. Provenance graphs encode these logs as interacting entities and events, exposing a causal and dependency structure that is often obscured in linear representations. Prior provenance-based detectors typically apply anomaly detection over such graphs, yet they frequently incur high false-positive rates and produce coarse grained alerts; moreover, approaches that heavily depend on node-specific identifiers (e.g., file paths) can learn spurious correlations, reducing robustness and limiting reliability across heterogeneous workloads. In this paper, we present Self-Training Adaptive Graph Encoder (stage), a lightweight, self-supervised anomaly detection framework for provenance graphs that (i) trains without attack labels and (ii) enforces leakage-free model selection and thresholding with explicit control over false-alarm rates. STAGE uses learnable degree and node-type embeddings, processed by a compact two-layer Graph Convolutional Networks (GCN) with residual connections and dual pooling. A memory augmented attention module captures global benign prototypes, improving resilience to rare-but-legitimate behaviors, and suppressing false alarms. Training combines contrastive learning over augmented graph views with a one-class Support Vector Data Description (SVDD) objective that learns a compact benign hypersphere in the embedding space. Inference, STAGE fuses neural embeddings with fixed dimensional structural graph statistics and scores them using an ensemble of classical one-class detectors. As a result, STAGE attains strong ranking quality and practical operating points on two benchmarks: the StreamSpot and Wget datasets. In the StreamSpot dataset, STAGE achieves an AUC of 0.998, operating at 95% recall with a 0% false positive rate. On the Wget dataset, it attains an AUC of 0.998 and an average precision of 0.998, achieving 100% recall and 96% precision at a 4% false positive rate. Overall, STAGE demonstrates strong empirical separability for benign-only provenance-based detection and provides an explicit mechanism to trade off recall and false positive rate through predefined thresholding policies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Marwan Ali Albahar (2026) studied this question.

synapsesocial.com/papers/69f154c0879cb923c4944f77https://doi.org/10.32604/cmc.2026.079941
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Provenance Graph-Based Deep Learning Framework for APT Detection in Edge Computing2025 · 3 citations
  2. 2Provenance Graph Modeling and Feature Enhancement for Power System APT Detection2025 · 2 citations
  3. 3Detecting advanced persistent threats via heterogeneous graph learning from homophily and heterogeneity views2026
  4. 4P3GNN: A Privacy-Preserving Provenance Graph-Based Model for APT Detection in Software Defined Networking2024 · 2 citations
  5. 5ProcSAGE: an efficient host threat detection method based on graph representation learning2024 · 10 citations