PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 15, 2026Journal of Data and Information Science0 citationsOpen Access

Scoring Structured Academic Documents with Large Language Models: Impact Case Studies

View Full Paper
MTMike ThelwallKKKayvan KoushaGHGuoxiu He

Key Points

  • This research investigates the effectiveness of medium-sized large language models in scoring structured academic documents like Impact Case Studies.
  • Evaluated five popular medium-sized LLMs on 6,010 REF 2021 ICSs.
  • Correlated estimated scores from LLMs with a proxy quality rating based on departmental average scores.
  • Tested scoring efficacy using different sections of ICSs to optimize score predictions.
  • Moderate correlation with proxy quality rating observed (highest Spearman correlation of 0.37).
  • Most LLMs performed statistically better than random guessing at section-level scoring.
  • A combined approach focusing on the summary and impact detail sections yielded the best scoring performance.

Abstract

Abstract Purpose Academic documents require expert time to evaluate, and Large Language Models (LLMs) might support this through score or decision predictions. For confidential structured academic texts, such as grants and Impact Case Studies (ICSs), medium-sized LLMs can be run offline without expensive computing infrastructures, enhancing security. Design/methodology/approach This study evaluates for the first time how well medium-sized LLMs can score structured academic documents using the UK Research Excellence Framework (REF) 2021 ICSs, and whether LLMs can guess scores from individual sections. We obtained score estimates from five recent popular LLMs (DeepSeek R1 32B, Qwen 3 32B, Magistral Small 24B, Gemma 3 27B, and Llama 4 Scout 27B) across 6,010 REF 2021 ICSs, correlating the scores with a proxy quality rating (departmental average score). Findings Scoring the full texts was only moderately effective (in terms of correlations with the proxy quality rating) and Llama 4 failed to score most of the longest. Surprisingly, all LLMs except Magistral were able to make statistically significantly above random guesses at ICS scores from each of the individual component sections (summary, underpinning research, references, details of the impacts, and sources to support the impact). A logical two-stage approach mimicking the human reviewer instructions did not outperform focusing on impact alone. The best strategy was to score the summary and the details of the impact sections combined (five times, averaged) with Gemma 3. This gave the highest Spearman correlation (0.37) with departmental average proxy quality scores (0.55 for department-level correlations). Practical implications Medium sized LLMs can be used to score structured academic documents to support research assessments. Research limitations This uses a single large case study with a public, albeit obscured, gold standard. Originality/value This improves on the state of the art despite the additional restrictions and with a much cheaper and potentially private open weights LLM approach.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Thelwall et al. (2026) studied this question.

synapsesocial.com/papers/6a2f9718a1cfeec490828337https://doi.org/10.1515/jdis-2025-0465
Ask AI
Helpful
Bookmark
Share
View Full Paper