Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 8, 2025Open Access

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models

View Full Paper
Ask AI
Bookmark
Share

Authors

HSHaoran SunZZZekun ZhangSZShaoning Zeng

Discussion

Loading...

Member takes

Overview

This approach improves alignment in large language models by addressing uncertainty in semantics, factuality, and value alignment.

Key Points

  • UDASA significantly improves the alignment of large language models with human intent and safety norms.
  • The framework offers automated alignment using multiple response generation and quantifying uncertainty across three dimensions.
  • Training samples are categorized into conservative, moderate, and exploratory stages based on their uncertainty difference.
  • Results indicate that UDASA outperforms existing methods in harmlessness, helpfulness, truthfulness, and sentiment generation.

Cite This Study

Sun et al. (2025) studied this question.

synapsesocial.com/papers/68e6860af44b9035634c209ehttps://doi.org/10.48550/arxiv.2507.17477
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Uncertainty Aware Learning for Language Model Alignment2024 · 1 citations
  2. 2ABC Align: Large Language Model Alignment for Safety & Accuracy2024
  3. 3Understanding the Learning Dynamics of Alignment with Human Feedback2024
  4. 4Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations2024
  5. 5Improving the Robustness of Large Language Models via Consistency Alignment2024 · 4 citations