PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 8, 2026Transactions of the Association for Computational Linguistics0 citationsOpen Access

Aligned Probing: Relating Toxic Behavior and Model Internals

View Full Paper
AWAndreas WaldisVGVagrant GautamALAnne Lauscher

Key Points

  • The aim is to explore how language models encode toxicity information in their outputs and internal structures.
  • Introduced aligned probing as a new interpretability framework.
  • Examined over 20 models including OLMo, Llama, and Mistral.
  • Conducted four case studies on detoxification, model quantization, and pre-training dynamics.
  • Language models encode toxicity information strongly, especially in lower layers.
  • Models generate less toxic outputs when they recognize input toxicity.
  • Model behavior varies significantly across different attributes like 'Threat'.

Abstract

Abstract Warning: This paper contains offensive text. We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity. alignedprobing.github.io

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Waldis et al. (2026) studied this question.

synapsesocial.com/papers/69d5f07d74eaea4b11a79f6bhttps://doi.org/10.1162/tacl.a.613
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Aya 23: Open Weight Releases to Further Multilingual Progress2024 · 9 citations
  2. 2Preference Tuning For Toxicity Mitigation Generalizes Across Languages2024 · 2 citations
  3. 3State of What Art? A Call for Multi-Prompt LLM Evaluation2024 · 99 citations
  4. 4Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization2024 · 1 citations
  5. 5From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP2024 · 2 citations