PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 20, 2025Machine Learning and Knowledge Extraction6 citationsOpen Access

Leveraging LLMs for Automated Extraction and Structuring of Educational Concepts and Relationships

View Full Paper
TYTong Qing YangBRBaofeng RenCGChenghao Gu

Key Points

  • LLMs have the potential to automate the extraction of educational concepts, improving the efficiency of course recommendations.
  • GPT-3.5 recorded the highest scores in quantitative metrics, but GPT-4o models produced more meaningful concepts overall.
  • Performance was assessed through automated experiments and human evaluations, illustrating how prompt design affects results.
  • Despite promising outcomes, LLM outputs still require expert revisions, indicating a need for careful implementation.

Abstract

Students must navigate large catalogs of courses and make appropriate enrollment decisions in many online learning environments. In this context, identifying key concepts and their relationships is essential for understanding course content and informing course recommendations. However, identifying and extracting concepts can be an extremely labor-intensive and time-consuming task when it has to be done manually. Traditional NLP-based methods to extract relevant concepts from courses heavily rely on resource-intensive preparation of detailed course materials, thereby failing to minimize labor. As recent advances in large language models (LLMs) offer a promising alternative for automating concept identification and relationship inference, we thoroughly investigate the potential of LLMs in automatically generating course concepts and their relations. Specifically, we systematically evaluate three LLM variants (GPT-3.5, GPT-4o-mini, and GPT-4o) across three distinct educational tasks, which are concept generation, concept extraction, and relation identification, using six systematically designed prompt configurations that range from minimal context (course title only) to rich context (course description, seed concepts, and subtitles). We systematically assess model performance through extensive automated experiments using standard metrics (Precision, Recall, F1, and Accuracy) and human evaluation by four domain experts, providing a comprehensive analysis of how prompt design and model choice influence the quality and reliability of the generated concepts and their interrelations. Our results show that GPT-3.5 achieves the highest scores on quantitative metrics, whereas GPT-4o and GPT-4o-mini often generate concepts that are more educationally meaningful despite lexical divergence from the ground truth. Nevertheless, LLM outputs still require expert revision, and performance is sensitive to prompt complexity. Overall, our experiments demonstrate the viability of LLMs as a tool for supporting educational content selection and delivery.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2025) studied this question.

synapsesocial.com/papers/68d46fc631b076d99fa69bc5https://doi.org/10.3390/make7030103
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Term-weighting approaches in automatic text retrieval1988 · 9,728 citations
  2. 2Boosting Course Recommendation Explainability: A Knowledge Entity Aware Model Using Deep Learning2024 · 3 citations
  3. 3Resources Sequencing Using Automatic Prerequisite--Outcome Annotation2015 · 25 citations
  4. 4Large Language Models for Recommendation: Progresses and Future Directions2023 · 30 citations
  5. 5A survey on large language models for recommendation2024 · 431 citations