PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Drones3 citationsOpen Access

Securing UAV Swarms with Vision Transformers: A Byzantine-Robust Federated Learning Framework for Cross-Modal Intrusion Detection

View Full Paper
CŞCanan Batur Şahin

Key Points

  • The study aims to develop a robust and privacy-preserving framework for UAV intrusion detection using federated learning.
  • Developed CM-BRF-ViT, integrating Vision Transformer architecture with Gramian Angular Field transformations.
  • Implemented Reference-GAF Consistency Aggregation to ensure model consistency against Byzantine attacks.
  • Conducted extensive experiments on UAVIDS-2025 and Cyber-Physical datasets to evaluate performance.
  • Utilized a learnable cross-modal fusion head for enhanced threat detection across different data types.
  • Achieved 97.1% detection accuracy on UAV network traffic and 78.5% on cyber-physical data.
  • Obtained a fused detection AUC of 0.993, indicating strong detection performance.
  • Maintained 89.6% accuracy under 40% Byzantine attacks, outperforming FedAvg by over 44 percentage points.

Abstract

The increasing deployment of uncrewed aerial vehicles (UAVs) in cyber-physical and safety-critical missions has amplified the need for intrusion detection systems that are accurate, privacy-preserving, and resilient to adversarial manipulation. In this paper, we propose CM-BRF-ViT, a Cross-Modal Byzantine-Robust Federated Vision Transformer framework for UAV intrusion detection that jointly addresses heterogeneous attack modeling, distributed learning security, and adaptive decision fusion. The proposed framework integrates Gramian Angular Field (GAF) transformations with Vision Transformer (ViT) architectures to effectively convert tabular network and cyber-physical features into discriminative visual representations suitable for attention-based learning. To enable privacy-preserving collaboration across distributed UAV nodes, CM-BRF-ViT operates within a federated learning paradigm and introduces Reference-GAF Consistency Aggregation (ReGCA). This novel Byzantine-robust aggregation mechanism jointly measures prediction consistency and feature-level semantic consistency using a trusted reference set and MAD-based robust weighting. Unlike conventional defenses that rely solely on parameter-space filtering, ReGCA supervises model updates at both behavioral and representation levels, significantly enhancing robustness against malicious clients. In addition, a learnable cross-modal fusion head is developed to adaptively combine attack probabilities derived from cyber and cyber-physical modalities, allowing the framework to exploit complementary threat signatures across layers. Extensive experiments conducted on the UAVIDS-2025 and Cyber-Physical datasets demonstrate that the proposed method achieves 97.1% detection accuracy for UAV network traffic and 78.5% for cyber-physical data, with a fused detection AUC of 0.993. Under adversarial settings, CM-BRF-ViT preserves 89.6% accuracy with up to 40% Byzantine clients, outperforming FedAvg by more than 44 percentage points. Ablation studies further confirm that ReGCA, cross-modal fusion, and ViT-based representation learning contribute complementary performance gains over baseline federated and centralized approaches. These results demonstrate that CM-BRF-ViT provides a robust, adaptive, and privacy-aware intrusion detection solution for UAV systems, making it well-suited for deployment in adversarial and resource-constrained aerial networks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Canan Batur Şahin (2026) studied this question.

synapsesocial.com/papers/698d6eeb5be6419ac0d54e26https://doi.org/10.3390/drones10020125
Ask AI
Helpful
Bookmark
Share
View Full Paper