PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving

View Full Paper
RWRuida WangBeijing Institute of Fashion TechnologyRPRui PanUniversity of Science and Technology of ChinaYLYuxin LiNorth Sichuan Medical University

Key Points

  • A new framework improves accuracy in formal theorem proving using Lean4, achieving 61.07%.
  • Extensive experiments show MA-LoT significantly outperforms previous methods, enhancing proof generation capabilities.
  • The model separates natural language tasks for proof generation and error analysis using a collaborative approach.
  • Combining long chain-of-thought reasoning with formal verification offers insightful generation potentials.

Abstract

Solving mathematical problems using computer-verifiable languages like Lean has significantly impacted the mathematical and computer science communities. State-of-the-art methods utilize a single Large Language Model (LLM) to generate complete proof or perform tree search, but they fail to balance these tasks. We propose **MA-LoT**: *Model-CollAboration Lean-based Long Chain-of-Thought*, a comprehensive framework for Lean4 theorem proving to solve this issue. It separates the cognition tasks of general NL for whole-proof generation and error analysis for proof correction using the model-collaboration method. We achieve this by structured interaction of the LLM and Lean4 verifier in Long CoT. To implement the framework, we propose the novel *LoT-Transfer Learning* training-inference pipeline, which enables the Long CoT thinking capability to LLMs without special data annotation. Extensive experiment shows that our framework achieves a **61.07%** accuracy rate on the Lean4 version of the MiniF2F-Test dataset, largely outperforming DeepSeek-V3 (33.61%), single-model tree search (InternLM-Step-Prover, 50.70%), and whole-proof generation (Godel-Prover, 55.33%) baselines. Furthermore, our findings highlight the potential of combining Long CoT with formal verification for a more insightful generation in a broader perspective.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68da58d1c1728099cfd10e58https://doi.org/10.48550/arxiv.2503.03205
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts2024
  2. 2LeanReasoner: Boosting Complex Logical Reasoning with Lean2024
  3. 3DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data2024 · 6 citations
  4. 4Lean Copilot: Large Language Models as Copilots for Theorem Proving in Lean2024 · 4 citations
  5. 5Lean-STaR: Learning to Interleave Thinking and Proving2024 · 2 citations