Key points are not available for this paper at this time.
Automated linkage between geoscientific literature and datasets is essential for improving data reuse, reproducibility, and knowledge discovery, yet existing methods often struggle with implicit dataset references, heterogeneous spatial–temporal expressions, and inconsistent naming conventions. To address this problem, we propose a literature–data linkage framework that integrates candidate retrieval, large language model (LLM)-based structured extraction, normalization, and knowledge graph construction. The framework first identifies candidate fragments through BM25-based retrieval, regex filtering, and whitelist-assisted scoring, and then applies schema-constrained prompting to extract dataset names and key attributes, including temporal coverage, spatial scope, resolution, provider, and role. The extracted results are subsequently normalized to canonical forms and ingested into a Neo4j-based knowledge graph linking articles, datasets, institutions, and regions. Experiments on a cross-journal benchmark show that the proposed framework achieves 93.79% precision, 90.66% recall, and 92.20% F1-score. Comparative experiments across multiple LLM backbones further indicate that the framework remains effective across both proprietary and open-source models, while ablation results confirm that candidate retrieval and normalization are the two most influential components for balanced extraction performance. The resulting knowledge graph provides a structured representation of literature–data linkages and supports exploration of dataset reuse patterns, provenance relations, and cross-document connections. These results demonstrate that carefully constrained LLM extraction, combined with retrieval and normalization, provides a robust and interpretable pathway for transforming unstructured geoscientific literature into structured and reusable knowledge.
Chen et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: