PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 6, 20260 citationsOpen Access

Systematically Identifying, Defining and Organizing Knowledge Components for Data Science Problem Solving through Human-LLM Collaboration

View Full Paper
FRFnu Priyanka RaniMAMaryam AlomairSPShimei Pan

Key Points

  • The study aims to identify and organize knowledge components essential for effective problem solving in data science.
  • Develop a framework combining human expertise and large language models (LLMs) for knowledge component identification.
  • Prompt multiple LLMs to generate decision points relevant to data science problem solving.
  • Synthesize and refine knowledge component definitions through human review and iteration.
  • Use sentence-embedding models to structure the resulting taxonomy.
  • Created a novel taxonomy of knowledge components for data science problem solving.
  • Demonstrated that LLMs can aid in generating and organizing domain knowledge for educational purposes.
  • Provided a scalable method for improving curriculum development and assessment tools in data science education.

Abstract

As demand grows for job-ready data science professionals, there is increasing recognition that traditional training often falls short in cultivating the higher-order reasoning and real-world problem-solving skills essential to the field. A foundational step toward addressing this gap is the identification and organization of knowledge components (KCs) that underlie data science problem solving (DSPS). KCs represent conditional knowledge-knowing about appropriate actions given particular contexts or conditions-and correspond to the critical decisions data scientists must make throughout the problem-solving process. While existing taxonomies in data science education support curriculum development, they often lack the granularity and focus needed to support the assessment and development of DSPS skills. In this paper, we present a novel framework that combines the strengths of large language models (LLMs) and human expertise to identify, define, and organize KCs specific to DSPS. We treat LLMs as ''knowledge engineering assistants'' capable of generating candidate KCs by drawing on their extensive training data, which includes a vast amount of domain knowledge and diverse sets of real-world DSPS cases. Our process involves prompting multiple LLMs to generate decision points, synthesizing and refining KC definitions across models, and using sentence-embedding models to infer the underlying structure of the resulting taxonomy. Human experts then review and iteratively refine the taxonomy to ensure validity. This human-AI collaborative workflow offers a scalable and efficient proof-of-concept for LLM-assisted knowledge engineering. The resulting KC taxonomy lays the groundwork for developing fine-grained assessment tools and adaptive learning systems that support deliberate practice in DSPS. Furthermore, the framework illustrates the potential of LLMs not just as content generators but as partners in structuring domain knowledge to inform instructional design. Future work will involve extending the framework by generating a directed graph of KCs based on their input-output dependencies and validating the taxonomy through expert consensus and learner studies. This approach contributes to both the practical advancement of DSPS coaching in data science education and the broader methodological toolkit for AI-supported knowledge engineering.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rani et al. (2025) studied this question.

synapsesocial.com/papers/698585db8f7c464f230099bbhttps://doi.org/10.13016/m21kh8-izfy
Ask AI
Helpful
Bookmark
Share
View Full Paper