PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 24, 2026Artificial Intelligence Review0 citationsOpen Access

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

MNMihai Dan Nadas

Key Points

  • To provide a comprehensive overview of advances, methodologies, and limitations in using large language models as automated judges for text and code generation from 2020 through early 2026.
  • Synthesized literature on LLM judging tasks across text domains (summarization, dialogue, factuality, safety) and code domains (correctness, code review, security).
  • Systematically reviewed evaluation prompting techniques (zero/few-shot, rubrics, pairwise comparison, chain-of-thought) and advanced multi-agent and tool-augmented pipelines.
  • Formulated a unified taxonomy for text and code judging, with core applications in benchmark evaluation, data filtering, and reward modeling for reinforcement learning.
  • Identified critical systematic vulnerabilities in automated judges, primarily length bias, position bias, and self-preference.
  • Highlighted calibration, fairness, reproducibility, and adversarial robustness as key open challenges requiring standardized protocols and uncertainty estimation.

Abstract

Abstract Large Language Models (LLMs) are increasingly used as judges to evaluate, rank, and critique AI-generated text and code. This survey provides a comprehensive overview of recent advances (2020–early 2026) in LLM-based evaluation, covering techniques, applications, and challenges across domains. We make three main contributions: (1) a unified taxonomy of LLM judging tasks spanning text (summarization, dialogue, factuality, safety) and code (correctness checking, code review, security analysis); (2) a systematic review of prompting strategies (zero/few-shot, rubric-based, pairwise comparison, chain-of-thought) and advanced pipelines (ensemble judges, multi-agent debate, tool-augmented verification); and (3) an analysis of LLM judge quality, documenting systematic biases (length, position, self-preference) and their mitigations. We review practical applications including benchmark evaluation (MT-Bench, Chatbot Arena), data filtering, and reward modeling for RLHF/RLAIF. Key challenges discussed include calibration, fairness, reproducibility, and adversarial robustness. We conclude with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration. LLM-based judging shows promise for scalable evaluation, but careful design and rigorous validation are essential to ensure these AI judges meet human standards of accuracy and fairness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mihai Dan Nadas (2026) studied this question.

synapsesocial.com/papers/6a8c0086bca056c88e6df5b0https://doi.org/10.1007/s10462-026-11652-0
Ask AI
Helpful
Bookmark
Share
View Full Paper