Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
April 9, 2024Open Access

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

View Full Paper
Ask AI
Bookmark
Share

Authors

KFKehua FengZhejiang University of Science and TechnologyKDKeyan DingCity University of Hong KongKMKede MaCity University of Hong Kong

Discussion

Loading...

Member takes

Implication

Key Points

Key points are not available for this paper at this time.

Cite This Study

Feng et al. (2024) studied this question.

synapsesocial.com/papers/68e6febab6db6435876790eahttps://doi.org/10.48550/arxiv.2404.08008
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Diagnosing Bias and Instability in LLM Evaluation: A Scalable Pairwise Meta-Evaluator2025 · 14 citations
  2. 2League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models2026
  3. 3What is the best model? Application-driven Evaluation for Large Language Models2024 · 1 citations
  4. 4Large Language Models are Inconsistent and Biased Evaluators2024 · 13 citations
  5. 5Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators2024 · 10 citations