PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 20255 citationsOpen Access

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

View Full Paper
SLSongyang LiuCLChaozhuo LiJQJianhua Qiu

Key Points

  • The survey identifies significant safety risks in large language models, including issues of toxicity and bias.
  • Safety evaluation encompasses distinct tasks such as assessing ethics, robustness, and truthfulness of LLM outputs.
  • Recent advancements in natural language processing necessitate a robust framework for evaluating safety metrics and benchmarks.
  • Prioritizing safety evaluation is essential for the responsible deployment of large language models in practical applications.

Abstract

With the rapid advancement of artificial intelligence technology, Large Language Models (LLMs) have demonstrated remarkable potential in the field of Natural Language Processing (NLP), including areas such as content generation, human-computer interaction, machine translation, and code generation, among others. However, their widespread deployment has also raised significant safety concerns. In recent years, LLM-generated content has occasionally exhibited unsafe elements like toxicity and bias, particularly in adversarial scenarios, which has garnered extensive attention from both academia and industry. While numerous efforts have been made to evaluate the safety risks associated with LLMs, there remains a lack of systematic reviews summarizing these research endeavors. This survey aims to provide a comprehensive and systematic overview of recent advancements in LLMs safety evaluation, focusing on several key aspects: (1) "Why evaluate" that explores the background of LLMs safety evaluation, how they differ from general LLMs evaluation, and the significance of such evaluation; (2) "What to evaluate" that examines and categorizes existing safety evaluation tasks based on key capabilities, including dimensions such as toxicity, robustness, ethics, bias and fairness, truthfulness, and so on; (3) "Where to evaluate" that summarizes the evaluation metrics, datasets and benchmarks currently used in safety evaluations; (4) "How to evaluate" that reviews existing evaluation toolkit, and categorizing mainstream evaluation methods based on the roles of the evaluators. Finally, we identify the challenges in LLMs safety evaluation and propose potential research directions to promote further advancement in this field. We emphasize the importance of prioritizing LLMs safety evaluation to ensure the safe deployment of these models in real-world applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2025) studied this question.

synapsesocial.com/papers/68e5c1c76950a706b22b5de5https://doi.org/10.48550/arxiv.2506.11094
Ask AI
Helpful
Bookmark
Share
View Full Paper