PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 5, 20240 citationsOpen Access

Revisiting a Pain in the Neck: Semantic Phrase Processing Benchmark for Language Models

View Full Paper
YLYang LiuMQMelissa Xiaohui QinHLHongming Li

Key Points

Key points are not available for this paper at this time.

Abstract

We introduce LexBench, a comprehensive evaluation suite enabled to test language models (LMs) on ten semantic phrase processing tasks. Unlike prior studies, it is the first work to propose a framework from the comparative perspective to model the general semantic phrase (i. e. , lexical collocation) and three fine-grained semantic phrases, including idiomatic expression, noun compound, and verbal construction. Thanks to, we assess the performance of 15 LMs across model architectures and parameter scales in classification, extraction, and interpretation tasks. Through the experiments, we first validate the scaling law and find that, as expected, large models excel better than the smaller ones in most tasks. Second, we investigate further through the scaling semantic relation categorization and find that few-shot LMs still lag behind vanilla fine-tuned models in the task. Third, through human evaluation, we find that the performance of strong models is comparable to the human level regarding semantic phrase processing. Our benchmarking findings can serve future research aiming to improve the generic capability of LMs on semantic phrase comprehension. Our source code and data are available at https: //github. com/jacklanda/LexBench

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2024) studied this question.

synapsesocial.com/papers/68e6b802b6db643587639507https://doi.org/10.48550/arxiv.2405.02861
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models2024
  2. 2The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models2024
  3. 3CogBench: a large language model walks into a psychology lab2024 · 5 citations
  4. 4LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs2025
  5. 5Large Language Model Benchmarks: A Taxonomy of Capabilities, Scientific Quality Assessment, and Saturation Analysis2026