PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 21, 20241 citationsOpen Access

Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models

View Full Paper
COCharles O’NeillTBThang Bui

Key Points

Key points are not available for this paper at this time.

Abstract

This paper introduces an efficient and robust method for discovering interpretable circuits in large language models using discrete sparse autoencoders. Our approach addresses key limitations of existing techniques, namely computational complexity and sensitivity to hyperparameters. We propose training sparse autoencoders on carefully designed positive and negative examples, where the model can only correctly predict the next token for the positive examples. We hypothesise that learned representations of attention head outputs will signal when a head is engaged in specific computations. By discretising the learned representations into integer codes and measuring the overlap between codes unique to positive examples for each head, we enable direct identification of attention heads involved in circuits without the need for expensive ablations or architectural modifications. On three well-studied tasks - indirect object identification, greater-than comparisons, and docstring completion - the proposed method achieves higher precision and recall in recovering ground-truth circuits compared to state-of-the-art baselines, while reducing runtime from hours to seconds. Notably, we require only 5-10 text examples for each task to learn robust representations. Our findings highlight the promise of discrete sparse autoencoders for scalable and efficient mechanistic interpretability, offering a new direction for analysing the inner workings of large language models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

O’Neill et al. (2024) studied this question.

synapsesocial.com/papers/68e69359b6db643587619d49https://doi.org/10.48550/arxiv.2405.12522
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models2024 · 4 citations
  2. 2Automatically Identifying Local and Global Circuits with Linear Computation Graphs2024 · 2 citations
  3. 3Transcoders Find Interpretable LLM Feature Circuits2024 · 6 citations
  4. 4Interpreting Attention Layer Outputs with Sparse Autoencoders2024 · 4 citations
  5. 5Scaling and evaluating sparse autoencoders2024 · 10 citations