PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2016Database904 citationsOpen Access

BioCreative V CDR task corpus: a resource for chemical disease relation extraction

View Full Paper
JLJiao LiYSYueping SunRJRobin J. Johnson

Key Points

  • This research aims to develop a comprehensive corpus for disease and chemical relations to support biomedical text-mining tasks.
  • Created the BC5CDR corpus with 1500 PubMed articles containing disease and chemical annotations.
  • Employed Medical Subject Headings (MeSH) indexers and CTD curators for entity annotation and relation extraction.
  • Ensured high quality through detailed guidelines and automatic tools; assessed inter-annotator agreement.
  • The corpus includes 4409 annotated chemicals and 5818 diseases with 3116 chemical-disease interactions.
  • Achieved inter-annotator agreement scores of 87.49% for diseases and 96.05% for chemicals using Jaccard similarity.
  • Successfully applied the BC5CDR corpus in BioCreative V challenge tasks.

Abstract

Community-run, formal evaluations and manually annotated text corpora are critically important for advancing biomedical text-mining research. Recently in BioCreative V, a new challenge was organized for the tasks of disease named entity recognition (DNER) and chemical-induced disease (CID) relation extraction. Given the nature of both tasks, a test collection is required to contain both disease/chemical annotations and relation annotations in the same set of articles. Despite previous efforts in biomedical corpus construction, none was found to be sufficient for the task. Thus, we developed our own corpus called BC5CDR during the challenge by inviting a team of Medical Subject Headings (MeSH) indexers for disease/chemical entity annotation and Comparative Toxicogenomics Database (CTD) curators for CID relation annotation. To ensure high annotation quality and productivity, detailed annotation guidelines and automatic annotation tools were provided. The resulting BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions. Each entity annotation includes both the mention text spans and normalized concept identifiers, using MeSH as the controlled vocabulary. To ensure accuracy, the entities were first captured independently by two annotators followed by a consensus annotation: The average inter-annotator agreement (IAA) scores were 87.49% and 96.05% for the disease and chemicals, respectively, in the test set according to the Jaccard similarity coefficient. Our corpus was successfully used for the BioCreative V challenge tasks and should serve as a valuable resource for the text-mining research community.Database URL: http://www.biocreative.org/tasks/biocreative-v/track-3-cdr/.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2016) studied this question.

synapsesocial.com/papers/69de9459741e97d2d4e93e79https://doi.org/10.1093/database/baw068
Ask AI
Helpful
Bookmark
Share
View Full Paper