PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 24, 20242 citationsOpen Access

MLPs Learn In-Context

View Full Paper
WTWilliam L. TongCPCengiz Pehlevan

Key Points

  • MLPs demonstrate in-context learning capabilities similar to advanced transformer models, enhancing their adaptability.
  • On certain relational reasoning tasks, MLPs outperform Transformers while operating within the same compute budget.
  • Multi-layer perceptrons and MLP-Mixer models show competitive performance against transformers in learning tasks designed to test relations directly in context—demonstrating versatility in machine learning architectures regardless of inductive bias pursuit. Lacking dependence on attention mechanisms, MLPs challenge traditional views on how learning tasks can be approached effectively—in light of new empirical evidence supporting the efficacy of simpler methods.

Abstract

In-context learning (ICL), the remarkable ability to solve a task from only input exemplars, has commonly been assumed to be a unique hallmark of Transformer models. In this study, we demonstrate that multi-layer perceptrons (MLPs) can also learn in-context. Moreover, we find that MLPs, and the closely related MLP-Mixer models, learn in-context competitively with Transformers given the same compute budget. We further show that MLPs outperform Transformers on a subset of ICL tasks designed to test relational reasoning. These results suggest that in-context learning is not exclusive to Transformers and highlight the potential of exploring this phenomenon beyond attention-based architectures. In addition, MLPs' surprising success on relational tasks challenges prior assumptions about simple connectionist models. Altogether, our results endorse the broad trend that ``less inductive bias is better" and contribute to the growing interest in all-MLP alternatives to task-specific architectures.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tong et al. (2024) studied this question.

synapsesocial.com/papers/68e6886ab6db643587610f3chttps://doi.org/10.48550/arxiv.2405.15618
Ask AI
Helpful
Bookmark
Share
View Full Paper