This is the authors' abstract. We don't add key points for this paper.
The emergence of large language models (LLMs) has significantly advanced natural language understanding, generation, and reasoning across various domains, including the biomedical field. Despite these advancements, the evaluation of biomedical LLMs remains limited, primarily relying on manually crafted datasets which are insufficient for comprehensively assessing LLMs’ capabilities. To address this challenge and cater specifically to the requirements of biomedical LLMs, we propose KGMedQA, an innovative evaluation benchmark based on knowledge graphs (KGs) designed to assess the knowledge and reasoning abilities of LLMs. Through careful alignment between natural language and KG structures, KGMedQA can be applied to arbitrary KGs, enabling automated question/answer generation and LLM evaluation. By leveraging the advantages of KG structures, we design seven tasks of varying complexity and focus, accompanied by specialized evaluation metrics. Experiments conducted with KGMedQA involve ten different LLMs, including general and specialized biomedical models, tested across two KGs focusing on different types of biomedical knowledge and reasoning. Compared to traditional methods, our results uncover more novel insights. For instance, while specialized models exhibit strengths in knowledge, they have deficiencies in reasoning abilities compared to general models. Additionally, factors such as model scales and prompting methods also impact the performance of LLMs. Our benchmark represents advancements in the evaluation of domain-specific LLMs, offering an effective tool for future research and development. Our source code is available at: https://github.com/PerseidsMeteorShower/KGMedQA .
No takes yet. Share an insight, caveat, or question.
Hao et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: