PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 20250 citationsOpen Access

Rethinking and Benchmarking Large Language Models for Graph Reasoning

View Full Paper
YHYuwei HuUniversity of Wisconsin–MadisonXHXinyi HuangUniversity of LeedsZWZhewei WeiRenmin University of China

Key Points

  • Graph reasoning capabilities in large language models are significantly underestimated with existing benchmarking methods.
  • Simple-Reasoning-Then-Coding achieves nearly perfect accuracy on current benchmarks, outperforming GPT-4o-mini.
  • A new GraphAlgorithm benchmark has been constructed, consisting of 239 graph problems and over 3,000 test instances.
  • Redirecting reasoning focus from replicating to designing graph algorithms can greatly enhance performance.

Abstract

Large Language Models (LLMs) for Graph Reasoning have been extensively studied over the past two years, involving enabling LLMs to understand graph structures and reason on graphs to solve various graph problems, with graph algorithm problems being the most prevalent. Recent studies underscore the potential of LLMs in handling graph reasoning tasks, but their performance is underwhelming. In this work, we point out issues with existing methods and benchmarks, and rethink the direction that LLMs for graph reasoning should strive toward. We find that base models, e.g., GPT-4o-mini, are largely underestimated due to improper reasoning focus. Base models with reasoning focus redirected from replicating graph algorithms to designing them can easily solve most graph reasoning tasks in existing benchmarks. To truly evaluate the graph reasoning capabilities of LLMs, we construct a more challenging GraphAlgorithm benchmark, comprising 239 different graph problems and 3,041 test instances collected from 4 competition platforms. Finally, we introduce a simple and strong baseline Simple-Reasoning-Then-Coding (Simple-RTC)-which guides LLMs to design graph algorithms first and then code to address graph reasoning tasks. Simple-RTC achieves near-perfect accuracy on existing benchmarks and significantly outperforms GPT-4o-mini and all prior methods on the GraphAlgorithm benchmark. This strong baseline encourages further advancements in LLMs for Graph Reasoning in the future.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hu et al. (2025) studied this question.

synapsesocial.com/papers/68e25382d6d66a53c2474788https://doi.org/10.48550/arxiv.2509.24260
Ask AI
Helpful
Bookmark
Share
View Full Paper