PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 202329 citationsOpen Access

Emergent Linear Representations in World Models of Self-Supervised Sequence Models

NNNeel NandaALAndrew LeeMWMartin Wattenberg

Key Points

Key points are not available for this paper at this time.

Abstract

How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023a). In this work, we provide evidence of a closely related linear representation of the board. In particular, we show that probing for "my colour" vs. "opponent's colour" may be a simple yet powerful way to interpret the model's internal state. This precise understanding of the internal representations allows us to control the model's behaviour with simple vector arithmetic. Linear representations enable significant interpretability progress, which we demonstrate with further exploration of how the world model is computed.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nanda et al. (2023) studied this question.

synapsesocial.com/papers/6a1578519b87f33fc69f95b0https://doi.org/10.18653/v1/2023.blackboxnlp-1.2
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Discovering Latent Knowledge in Language Models Without Supervision2022 · 47 citations
  2. 2Inference-Time Intervention: Eliciting Truthful Answers from a Language Model2023 · 40 citations
  3. 3The Hydra Effect: Emergent Self-repair in Language Model Computations2023 · 3 citations
  4. 4Governance Architecture for Neural Network Superposition: A Structural Solution to Hallucination via Routing and Interference Filtering2022 · 48 citations
  5. 5Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023 · 79,073 citations