PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 20240 citationsOpen Access

Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

View Full Paper
DBDan BraunJTJordan TaylorNGNicholas Goldowsky-Dill

Key Points

Key points are not available for this paper at this time.

Abstract

Identifying the features learned by neural networks is a core challenge in mechanistic interpretability. Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary that reconstructs a network's internal activations, have been used to identify these features. However, SAEs may learn more about the structure of the datatset than the computational structure of the network. There is therefore only indirect reason to believe that the directions found in these dictionaries are functionally important to the network. We propose end-to-end (e2e) sparse dictionary learning, a method for training SAEs that ensures the features learned are functionally important by minimizing the KL divergence between the output distributions of the original model and the model with SAE activations inserted. Compared to standard SAEs, e2e SAEs offer a Pareto improvement: They explain more network performance, require fewer total features, and require fewer simultaneously active features per datapoint, all with no cost to interpretability. We explore geometric and qualitative differences between e2e SAE features and standard SAE features. E2e dictionary learning brings us closer to methods that can explain network behavior concisely and accurately. We release our library for training e2e SAEs and reproducing our analysis at https: //github. com/ApolloResearch/e2eₛae

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Braun et al. (2024) studied this question.

synapsesocial.com/papers/68e69aefb6db643587620783https://doi.org/10.48550/arxiv.2405.12241
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control2024 · 3 citations
  2. 2Improving Dictionary Learning with Gated Sparse Autoencoders2024 · 2 citations
  3. 3Analysis of Variational Sparse Autoencoders2025
  4. 4Disentangling Dense Embeddings with Sparse Autoencoders2024 · 3 citations
  5. 5A Comparative Analysis of Sparse Autoencoder and Activation Difference in Language Model Steering2025