We introduce IGLA (Inverse Geodesic Latent Attractor), a physics-inspired framework for detecting harmful intent in large language models (LLMs) using latent-space trajectory dynamics. Unlike traditional safety systems that rely on supervised classifiers or alignment training, IGLA operates entirely within the model’s hidden-state geometry without requiring labeled data or additional forward passes. The method models latent representations as particles evolving under an inverse-distance potential field defined over harm-category centroids. By simulating short Euler trajectories and analyzing convergence behavior, IGLA identifies harmful prompts based on their attraction toward unsafe latent regions. Evaluated on 787 prompts from AdvBench, HarmBench, and WildGuard, the system achieves 78.02% accuracy with only 5.50% inference overhead on consumer GPUs. Ablation studies show that trajectory dynamics contribute 100% of the performance gain over static proximity methods. We further report a novel geometric property of LLM safety data: synthetic augmentation reduces intra-basin variance, revealing a trade-off between learnability and adversarial robustness. This work serves as a proof-of-concept for physics-inspired latent-space safety analysis and demonstrates that meaningful safety signals exist in hidden-state dynamics without supervised training.
Nambiar et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: