PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 25, 2026Big Data and Cognitive Computing0 citationsOpen Access

Thinking Machines: Mathematical Reasoning in the Age of LLMs

View Full Paper
AAAndrea AspertiANAlberto NaiboCCClaudio Sacerdoti Coen

Key Points

  • The aim is to investigate the reasoning abilities of large language models in mathematical contexts.
  • Review of current state-of-the-art models in mathematical reasoning with LLMs.
  • Examination of benchmarks for evaluating LLMs in mathematical tasks.
  • Analysis of the trade-offs between traditional and formalized mathematics.
  • Identified significant challenges in proof synthesis compared to code generation.
  • Explored the role of feedback and supervision in enhancing reasoning.
  • Discussed the nature of logical states in LLMs and their implications for mathematical cognition.

Abstract

Large Language Models (LLMs) have demonstrated impressive capabilities in structured reasoning and symbolic tasks, with coding emerging as a particularly successful application. This progress has naturally motivated efforts to extend these models to mathematics, both in its traditional form, expressed through natural-style mathematical language, and in its formalized counterpart, expressed in a symbolic syntax suitable for automatic verification. Yet, despite apparent parallels between programming and proof construction, advances in formalized mathematics have proven significantly more challenging. This gap raises fundamental questions about the nature of reasoning in current LLM architectures, the role of supervision and feedback, and the extent to which such models maintain an internal notion of computational or deductive state. In this article, we review the current state-of-the-art in mathematical reasoning with LLMs, focusing on recent models and benchmarks. We explore three central issues at the intersection of machine learning and mathematical cognition: (i) the trade-offs between traditional and formalized mathematics as training and evaluation domains; (ii) the structural and methodological reasons why proof synthesis remains more brittle than code generation; and (iii) whether LLMs genuinely represent or merely emulate a notion of evolving logical state. Our goal is not to draw rigid distinctions but to clarify the present boundaries of these systems and outline promising directions for their extension.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Asperti et al. (2026) studied this question.

synapsesocial.com/papers/6975b4fd5a65d392b01e5cb8https://doi.org/10.3390/bdcc10010038
Ask AI
Helpful
Bookmark
Share
View Full Paper