PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 2020Journal of the American Medical Informatics Association85 citationsOpen Access

Potential limitations in COVID-19 machine learning due to data source variability: A case study in the nCov2019 dataset

View Full Paper
CSCarlos SáezNRNekane Romero-GarcíaJCJ. Alberto Conejero

Key Points

Key points are not available for this paper at this time.

Abstract

Abstract Objective The lack of representative coronavirus disease 2019 (COVID-19) data is a bottleneck for reliable and generalizable machine learning. Data sharing is insufficient without data quality, in which source variability plays an important role. We showcase and discuss potential biases from data source variability for COVID-19 machine learning. Materials and Methods We used the publicly available nCov2019 dataset, including patient-level data from several countries. We aimed to the discovery and classification of severity subgroups using symptoms and comorbidities. Results Cases from the 2 countries with the highest prevalence were divided into separate subgroups with distinct severity manifestations. This variability can reduce the representativeness of training data with respect the model target populations and increase model complexity at risk of overfitting. Conclusions Data source variability is a potential contributor to bias in distributed research networks. We call for systematic assessment and reporting of data source variability and data quality in COVID-19 data sharing, as key information for reliable and generalizable machine learning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sáez et al. (2020) studied this question.

synapsesocial.com/papers/6a0e9371a03ab94435045b23https://doi.org/10.1093/jamia/ocaa258
Ask AI
Helpful
Bookmark
Share
View Full Paper