PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 17, 20240 citationsOpen Access

Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs

View Full Paper
SWSiyuan WangZWZhongyu WeiYCYejin Choi

Key Points

  • LLMs show significant gaps in logic understanding compared to humans, especially with complex rules.
  • Evaluation of GPT models indicates limitations in their grasp of compositional and structural rules.
  • Our logic scaffolding framework, ULogic, generates an inferential rule base to enhance reasoning tasks effectively across various domains with 5 key areas explored with a multi-judger framework employed for improvement of reasoning performance. Such scaffolding may enable better performance on commonsense reasoning tasks, indicating a pathway for future model development.

Abstract

Large language models (LLMs) have achieved impressive human-like performance across various reasoning tasks. However, their mastery of underlying inferential rules still falls short of human capabilities. To investigate this, we propose a logic scaffolding inferential rule generation framework, to construct an inferential rule base, ULogic, comprising both primitive and compositional rules across five domains. Our analysis of GPT-series models over a rule subset reveals significant gaps in LLMs' logic understanding compared to human performance, especially in compositional and structural complex rules with certain bias patterns. We further distill these rules into a smaller-scale inference engine for flexible rule generation and enhancing downstream reasoning. Through a multi-judger evaluation, our inference engine proves effective in generating accurate, complex and abstract conclusions and premises, and improve various commonsense reasoning tasks. Overall, our work sheds light on LLMs' limitations in grasping inferential rule and suggests ways to enhance their logical reasoning abilities~Code and data are available at https: //github. com/SiyuanWangw/ULogic. .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2024) studied this question.

synapsesocial.com/papers/68e78cf2b6db6435876feb0ahttps://doi.org/10.48550/arxiv.2402.11442
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models2024 · 1 citations
  2. 2Do Large Language Models Understand Logic or Just Mimick Context?2024 · 3 citations
  3. 3LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning2025
  4. 4Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis2024
  5. 5Evidence of conceptual mastery in the application of rules by Large Language Models2025