PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 5, 20242 citationsOpen Access

Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models

View Full Paper
SWSheng-Lun WeiNational Taiwan UniversityCWCheng-Kuang WuHHHen‐Hsen HuangInstitute of Information Science, Academia Sinica

Key Points

  • Selection biases significantly affect the decision-making processes of large language models.
  • The empirical analysis evaluated multiple models across various tasks, indicating substantial sensitivity to option order and token usage.
  • Mitigation strategies were developed to enhance model performance, focusing on reducing biases associated with token and order sensitivity in LLMs. The findings can inform future applications and improvements in model reliability and robustness for selection tasks.

Abstract

In this paper, we investigate the phenomena of "selection biases" in Large Language Models (LLMs), focusing on problems where models are tasked with choosing the optimal option from an ordered sequence. We delve into biases related to option order and token usage, which significantly impact LLMs' decision-making processes. We also quantify the impact of these biases through an extensive empirical analysis across multiple models and tasks. Furthermore, we propose mitigation strategies to enhance model performance. Our key contributions are threefold: 1) Precisely quantifying the influence of option order and token on LLMs, 2) Developing strategies to mitigate the impact of token and order sensitivity to enhance robustness, and 3) Offering a detailed analysis of sensitivity across models and tasks, which informs the creation of more stable and reliable LLM applications for selection problems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wei et al. (2024) studied this question.

synapsesocial.com/papers/68e660e5b6db6435875ef370https://doi.org/10.48550/arxiv.2406.03009
Ask AI
Helpful
Bookmark
Share
View Full Paper