PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 20, 20260 citationsOpen Access

SAVANT: Sparse Adaptive Vector Alignment for Continual Learning and Persistent Agent Safety

View Full Paper
CSCleilson Elias SousaUniversidade Federal do Rio de Janeiro

Key Points

  • This research aims to address the challenges of alignment collapse during continual learning in large language models (LLMs).
  • Introduced CALMS-SAVANT architecture for persistent LLM agents
  • Implemented an asynchronous regulated plasticity approach
  • Evaluated performance using the Persistent Prompt Poisoning Benchmark over 100 epochs.
  • CALMS-SAVANT reduced Poison Success Rate from 89.7% to 4.8%
  • Fact Retention Score was maintained at 92.3%
  • Proven effective against adversarial pressures with enhanced stability.

Abstract

Official Preprint: CALMS-SAVANT Architecture Persistent LLM agents suffer from a critical vulnerability during continual learning: catastrophic consolidation. Under longitudinal adversarial pressure, standard replay mechanisms passively assimilate stealth poisoning, leading to permanent alignment collapse. This paper introduces CALMS-SAVANT, an asynchronously regulated plasticity architecture that solves this by replacing unrestricted parameter updates with rigorous geometric validation and Byzantine-robust initialization. Key Highlights: Multi-Seed Consensus Prior: Eliminates epoch-0 vulnerabilities by constructing the initial Identity Graph via Byzantine-fault-tolerant geometric arbitration. Geometric Stability Filter: Quantifies global manifold drift using Sinkhorn-regularized Wasserstein-1 distance. The Cognitive Diode: Executes adaptive gradient suppression to block adversarial parametric integration. PPP-Bench: Introduces the Persistent Prompt Poisoning Benchmark for longitudinal evaluation over 100 epochs. Results: Evaluated on LLaMA-3-8B, CALMS-SAVANT reduced the terminal Poison Success Rate (PSR) to 4.8% (down from 89.7% in standard replay), while safely preserving a Fact Retention Score (FRS) of 92.3%. Code and benchmark datasets are being prepared for release on GitHub.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cleilson Elias Sousa (2026) studied this question.

synapsesocial.com/papers/6a0d5064f03e14405aa9c2b5https://doi.org/10.5281/zenodo.20277576
Ask AI
Helpful
Bookmark
Share
View Full Paper