PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 12, 20241 citationsOpen Access

Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward

View Full Paper
XXXuan XieJSJiayang SongZZZhehua Zhou

Key Points

Key points are not available for this paper at this time.

Abstract

While Large Language Models (LLMs) have seen widespread applications across numerous fields, their limited interpretability poses concerns regarding their safe operations from multiple aspects, e.g., truthfulness, robustness, and fairness. Recent research has started developing quality assurance methods for LLMs, introducing techniques such as offline detector-based or uncertainty estimation methods. However, these approaches predominantly concentrate on post-generation analysis, leaving the online safety analysis for LLMs during the generation phase an unexplored area. To bridge this gap, we conduct in this work a comprehensive evaluation of the effectiveness of existing online safety analysis methods on LLMs. We begin with a pilot study that validates the feasibility of detecting unsafe outputs in the early generation process. Following this, we establish the first publicly available benchmark of online safety analysis for LLMs, including a broad spectrum of methods, models, tasks, datasets, and evaluation metrics. Utilizing this benchmark, we extensively analyze the performance of state-of-the-art online safety analysis methods on both open-source and closed-source LLMs. This analysis reveals the strengths and weaknesses of individual methods and offers valuable insights into selecting the most appropriate method based on specific application scenarios and task requirements. Furthermore, we also explore the potential of using hybridization methods, i.e., combining multiple methods to derive a collective safety conclusion, to enhance the efficacy of online safety analysis for LLMs. Our findings indicate a promising direction for the development of innovative and trustworthy quality assurance methodologies for LLMs, facilitating their reliable deployments across diverse domains.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xie et al. (2024) studied this question.

synapsesocial.com/papers/68e6f60eb6db6435876710f6https://doi.org/10.48550/arxiv.2404.08517
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs2025 · 5 citations
  2. 2ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming2024 · 3 citations
  3. 3S-Eval: Automatic and Adaptive Test Generation for Benchmarking Safety Evaluation of Large Language Models2024 · 1 citations
  4. 4LLM-Safety Evaluations Lack Robustness2025
  5. 5S-Eval: Towards Automated Safety Evaluation with Enhancement for Large Language Models2026