PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 20260 citationsOpen Access

Activation-Scaled ANN-to-SNN Conversion with SNN Guardrail: A Unified Framework for AI Interpretability, Hallucination Detection, and Real-Time Adversarial Defense

HFHiroto Funasaki

Key Points

  • To create a framework that enhances ANN-to-SNN conversion for efficiency, interpretability, and safety in AI systems.
  • Developed an SNN Guardrail for real-time AI safety and jailbreak detection.
  • Conducted scaling law discovery to analyze sensitivity with model size.
  • Validated frameworks with TinyLlama and ViT-Base architectures.
  • Performed analysis of black-box AI models using SNNs as computational microscopes.
  • Achieved 100% jailbreak detection rate across all attack types.
  • Demonstrated +3.1 and +4.2 increases in sensitivity for GPT-2 and TinyLlama respectively.
  • Observed significant deviations (+10 to +19σ TTFS) in neural instability detection during jailbreak attacks.
  • Maintained 100% accuracy preservation with a hybrid architecture.

Abstract

I present a unified framework that extends ANN-to-SNN conversion beyond efficiency optimization to enable novel AI interpretability analysis and real-time adversarial defense. My approach uses Spiking Neural Networks as "computational microscopes" to analyze black-box AI models. **v4 Updates:**- NEW: SNN Guardrail for real-time AI safety with **100% jailbreak detection rate** (8/8 attack types)- NEW: Scaling Law Discovery - TTFS sensitivity increases with model size (GPT-2: +3.1, TinyLlama 1.1B: +4.2)- NEW: TinyLlama (1.1B params) validation- NEW: Neural instability detection (jailbreak attacks cause +10 to +19σ TTFS deviation) **Previous Results (v1-v3):**- Universal threshold formula: θ = 2.0 × max(activation)- 100% accuracy preservation with hippocampal hybrid architecture- GPT-2 attention TTFS analysis: +3.1 increase for meaningless inputs- Hallucination detection: AUC 0.75 with ensemble classifier- ViT-Base (86M params) validation with CIFAR-100 Key insight: "Measure the AI's brainwaves, and block it when it's about to lie." Code: https://github.com/hafufu-stack/temporal-coding-simulation/tree/main/ann-to-snn-converter

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hiroto Funasaki (2026) studied this question.

synapsesocial.com/papers/6988292d0fc35cd7a8849404https://doi.org/10.5281/zenodo.18493943
Ask AI
Helpful
Bookmark
Share
View Full Paper