PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026China Scientific Data0 citationsOpen Access

A dataset of challenging mathematical problems for reasoning large language models: SD1K

DZDanhao ZhuFHFei Huang

Key Points

  • The aim is to create a robust dataset for training and evaluating large language models in mathematical reasoning.
  • Constructed SD1K dataset from 7,500 training samples of Sky-T1-32B-Preview model.
  • Extracted 9,054 mathematical concept entities with definitions and applications.
  • Generated new mathematical problems using random sampling of these entities.
  • Applied filtering with predefined rules and model-based verification to ensure quality.
  • Selected 1,000 high-quality and challenging mathematical problems.
  • Dataset supports the development of large language model performance in reasoning tasks.

Abstract

To address the lack of insufficient challenging training data for reasoning large language models, this study constructs a challenging mathematics dataset for reasoning, termed SD1K, using a methodology that integrates concept extraction and model synthesis. The dataset is built upon a collection of 7,500 training samples from the Sky-T1-32B-Preview model, from which 9,054 mathematical concept entities are extracted. Each concept entity includes a formal definition, application scenarios, and usage example. Using large language models, new mathematical problems are generated by randomly sampling these entities, with long-form reasoning chains and corresponding answers automatically constructed for each problem. A filtering process combining predefined rules and model-based verification is applied to select 1,000 high-quality and challenging mathematical problems, forming the final SD1K dataset. The SD1K dataset provides a valuable resource for the training and evaluation of large language models in reasoning tasks and contributes to the advancement of model performance in complex inference scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2026) studied this question.

synapsesocial.com/papers/69c8c3cede0f0f753b39ee46https://doi.org/10.11922/11-6035.csd.2025.0120.zh
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning2024
  2. 2Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models2025
  3. 3Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On2024
  4. 4Stepwise Self-Consistent Mathematical Reasoning with Large Language Models2024
  5. 5On The Main Challenges And New Perspectives On Generating Robust Benchmark Datasets For Large Language Models In Modern Advanced Mathematics2026