PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 20250 citationsOpen Access

Transformers Don't In-Context Learn Least Squares Regression

View Full Paper
JHJoshua HillBEBenjamin EyreECElliot Creager

Key Points

  • Transformers fail to generalize after distribution shifts in in-context learning tasks, indicating a limitation in their capabilities.
  • In contrast to ordinary least squares regression, transformers trained for in-context learning exhibit poor performance with out-of-distribution inputs.
  • The study employs a suite of out-of-distribution experiments to assess transformers' capacity for generalization through linear regression techniques.
  • A spectral analysis of the learned representations shows that the training corpus significantly impacts in-context learning behaviours.

Abstract

In-context learning (ICL) has emerged as a powerful capability of large pretrained transformers, enabling them to solve new tasks implicit in example input-output pairs without any gradient updates. Despite its practical success, the mechanisms underlying ICL remain largely mysterious. In this work we study synthetic linear regression to probe how transformers implement learning at inference time. Previous works have demonstrated that transformers match the performance of learning rules such as Ordinary Least Squares (OLS) regression or gradient descent and have suggested ICL is facilitated in transformers through the learned implementation of one of these techniques. In this work, we demonstrate through a suite of out-of-distribution generalization experiments that transformers trained for ICL fail to generalize after shifts in the prompt distribution, a behaviour that is inconsistent with the notion of transformers implementing algorithms such as OLS. Finally, we highlight the role of the pretraining corpus in shaping ICL behaviour through a spectral analysis of the learned representations in the residual stream. Inputs from the same distribution as the training data produce representations with a unique spectral signature: inputs from this distribution tend to have the same top two singular vectors. This spectral signature is not shared by out-of-distribution inputs, and a metric characterizing the presence of this signature is highly correlated with low loss.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hill et al. (2025) studied this question.

synapsesocial.com/papers/68de5da283cbc991d0a20800https://doi.org/10.48550/arxiv.2507.09440
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1In-Context Learning with Representations: Contextual Generalization of Trained Transformers2024
  2. 2Training Nonlinear Transformers for Efficient In-Context Learning: A Theoretical Learning and Generalization Analysis2024 · 1 citations
  3. 3Formalizing In-Context Learning in Transformers as Implicit Gradient Descent2026
  4. 4Learning Linear Regression with Low-Rank Tasks in-Context2025
  5. 5Does learning the right latent variables necessarily improve in-context learning?2024