PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 20, 2025Frontiers in Neuroinformatics8 citationsOpen Access

Large language models can extract metadata for annotation of human neuroimaging publications

View Full Paper
MTMatthew D. TurnerAAAbhishek AppajiNRNibras Ar Rakib

Key Points

  • The large language model achieved performance comparable to human annotators, showing a range of 0.91 to 0.97 on zero-shot prompts without feedback.
  • Recent evaluations confirmed that LLMs can effectively perform metadata extraction for neuroimaging publications with minimal manual effort.
  • Finding similarities between errors of LLMs and human annotators suggests that LLMs can be reliable in real-world annotation tasks.
  • The study encourages the development of specialized benchmarks for further improving the performance of LLMs in complex annotation scenarios.

Abstract

We show that recent (mid-to-late 2024) commercial large language models (LLMs) are capable of good quality metadata extraction and annotation with very little work on the part of investigators for several exemplar real-world annotation tasks in the neuroimaging literature. We investigated the GPT-4o LLM from OpenAI which performed comparably with several groups of specially trained and supervised human annotators. The LLM achieves similar performance to humans, between 0.91 and 0.97 on zero-shot prompts without feedback to the LLM. Reviewing the disagreements between LLM and gold standard human annotations we note that actual LLM errors are comparable to human errors in most cases, and in many cases these disagreements are not errors. Based on the specific types of annotations we tested, with exceptionally reviewed gold-standard correct values, the LLM performance is usable for metadata annotation at scale. We encourage other research groups to develop and make available more specialized “micro-benchmarks,” like the ones we provide here, for testing both LLMs, and more complex agent systems annotation performance in real-world metadata annotation tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Turner et al. (2025) studied this question.

synapsesocial.com/papers/68af4eaead7bf08b1ead7259https://doi.org/10.3389/fninf.2025.1609077
Ask AI
Helpful
Bookmark
Share
View Full Paper