PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models

View Full Paper
SSShashata SawmyaMAMicah AdlerNSNir Shavit

Key Points

  • Results show that interpretable categorical features emerge at distinct temporal and scale thresholds in large language models, challenging existing assumptions.
  • Spatial analysis uncovered unexpected reactivation of early-layer semantic features in later layers, indicating complex representational dynamics.
  • Mechanistic interpretability was achieved through the use of sparse autoencoders, allowing insight into the activation of semantic concepts.
  • Findings provide new understanding into the behavior of large language models at different training checkpoints and model scales.

Abstract

This paper studies the emergence of interpretable categorical features within large language models (LLMs), analyzing their behavior across training checkpoints (time), transformer layers (space), and varying model sizes (scale). Using sparse autoencoders for mechanistic interpretability, we identify when and where specific semantic concepts emerge within neural activations. Results indicate clear temporal and scale-specific thresholds for feature emergence across multiple domains. Notably, spatial analysis reveals unexpected semantic reactivation, with early-layer features re-emerging at later layers, challenging standard assumptions about representational dynamics in transformer models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sawmya et al. (2025) studied this question.

synapsesocial.com/papers/68da5a3ec1728099cfd11966https://doi.org/10.48550/arxiv.2505.19440
Ask AI
Helpful
Bookmark
Share
View Full Paper