PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 21, 20254 citations

Pre-Meta: Priors-augmented Retrieval for LLM-based Metadata Generation.

View Full Paper
PTPhil TinnSSSondre SørbøSJShanshan Jiang

Key Points

  • Pre-Meta improves metadata accuracy by 23% to 75% compared to standard methods across three LLMs.
  • Validation involved 1500 papers, focusing on five selected metadata fields to demonstrate the effectiveness of the method.
  • The approach uses an enriched retrieval procedure leveraging prior information for improved automated metadata generation.
  • This method highlights potential advances in data annotation processes, indicating effective tools for efficient genomic data handling.

Abstract

While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large Language Models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This paper presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tinn et al. (2025) studied this question.

synapsesocial.com/papers/68d46cbf31b076d99fa689c1https://doi.org/10.1093/bioinformatics/btaf519
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Enhanced semantic classification of microbiome sample origins using Large Language Models (LLMs)2026
  2. 2Enhanced semantic classification of microbiome sample origins using large language models (LLMs)2026 · 1 citations
  3. 3Identification of biomedical entities from multiple repositories using a specialized metadata schema and search-augmented large language models2026 · 2 citations
  4. 4Large language models can extract metadata for annotation of human neuroimaging publications2025 · 8 citations
  5. 5A new AI assisted approach aligns data standards and accelerates interoperability in biomedical research2026 · 2 citations