PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 19, 2026PLoS ONE0 citationsOpen Access

Pretraining effective T5 generative models for clinical and biomedical applications

View Full Paper
SASaad AlthabitiCCChuming ChenSASultan Alrowili

Key Points

  • The aim is to assess the effects of corpus selection and vocabulary design on T5 generative model performance in clinical and biomedical fields.
  • Developed five T5-EHR models pretrained on various clinical and biomedical corpora.
  • Conducted evaluations across multiple clinical and biomedical tasks to measure performance.
  • Analyzed the impact of different vocabulary choices on model outcomes.
  • Models pretrained on clinical data significantly outperform those with additional biomedical data for clinical tasks.
  • Clinical-specific vocabularies yield better performance than general biomedical vocabularies for tasks needing deep clinical comprehension.
  • T5 generative models show competitive performance against state-of-the-art discriminative models in biomedical applications.

Abstract

This paper presents a study of the impact of corpus selection and vocabulary design on the performance of T5-based language models in clinical and biomedical domains. We introduce five different T5-EHR models, each pretrained from scratch using different combinations of clinical and biomedical corpora alongside domain-specific vocabularies. We evaluated these models across a variety of clinical and biomedical tasks to quantify the impact of pretraining data and vocabulary tokenization choices on downstream performance. Our findings reveal the importance of aligning both pretraining corpus and vocabulary with the target domain. Models pretrained exclusively on clinical data achieve superior performance on clinical tasks, while adding biomedical data contributes only marginal gains in most cases, with a few exceptions. Similarly, the choice of vocabulary significantly influences model performance, with clinical-specific vocabularies outperforming general biomedical vocabularies in tasks requiring a deeper understanding of clinical language. Also, the T5 generative models perform competitively with state-of-the-art discriminative models on several biomedical benchmarks, demonstrating strong generalization to biomedical domain. Overall, these results emphasize that task-specific selection of corpus and vocabulary is essential for optimizing model performance in clinical and biomedical natural language processing (NLP).

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Althabiti et al. (2026) studied this question.

synapsesocial.com/papers/69e473bd010ef96374d8f7f5https://doi.org/10.1371/journal.pone.0342610
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1BioCreative V CDR task corpus: a resource for chemical disease relation extraction2016 · 904 citations
  2. 2MIMIC-IV, a freely accessible electronic health record dataset2023 · 3,198 citations
  3. 3The DDI corpus: An annotated corpus with pharmacological substances and drug–drug interactions2013 · 494 citations
  4. 42010 i2b2/VA challenge on concepts, assertions, and relations in clinical text2011 · 1,291 citations
  5. 5Deep learning for multiple sclerosis lesion classification and stratification using MRI2025 · 36 citations