PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 2017Sao Paulo Medical Journal78 citationsOpen Access

Comparison of machine-learning algorithms to build a predictive model for detecting undiagnosed diabetes - ELSA-Brasil: accuracy study

AOAndré Rodrigues OliveraVRValter RoeslerCICirano Iochpe

Key Result

Artificial neural networks and logistic regression models successfully identified individuals with undiagnosed diabetes, achieving mean areas under the curve of 74.17% and 74.41% respectively.

Study Design

Type

Observational

Structured PICO

Can machine-learning algorithms accurately detect undiagnosed diabetes using easily-obtained clinical data?

P
Population
Participants from the Longitudinal Study of Adult Health (ELSA-Brasil)
I
Intervention
Machine-learning algorithms (logistic regression, artificial neural network, naïve Bayes, K-nearest neighbor and random forest) using clinical data
C
Comparator
Comparison between different machine-learning algorithms
O
Outcome
Performance in detecting undiagnosed diabetes, measured by area under the curve (AUC)

Machine-learning models, particularly artificial neural networks and logistic regression, can feasibly identify individuals with a high probability of having undiagnosed diabetes using easily obtained clinical data.

Main Result

Absolute Event Rate: 74.17% vs 74.41%

Abstract

CONTEXT AND OBJECTIVE:: Type 2 diabetes is a chronic disease associated with a wide range of serious health complications that have a major impact on overall health. The aims here were to develop and validate predictive models for detecting undiagnosed diabetes using data from the Longitudinal Study of Adult Health (ELSA-Brasil) and to compare the performance of different machine-learning algorithms in this task. DESIGN AND SETTING:: Comparison of machine-learning algorithms to develop predictive models using data from ELSA-Brasil. METHODS:: After selecting a subset of 27 candidate variables from the literature, models were built and validated in four sequential steps: (i) parameter tuning with tenfold cross-validation, repeated three times; (ii) automatic variable selection using forward selection, a wrapper strategy with four different machine-learning algorithms and tenfold cross-validation (repeated three times), to evaluate each subset of variables; (iii) error estimation of model parameters with tenfold cross-validation, repeated ten times; and (iv) generalization testing on an independent dataset. The models were created with the following machine-learning algorithms: logistic regression, artificial neural network, naïve Bayes, K-nearest neighbor and random forest. RESULTS:: The best models were created using artificial neural networks and logistic regression. -These achieved mean areas under the curve of, respectively, 75.24% and 74.98% in the error estimation step and 74.17% and 74.41% in the generalization testing step. CONCLUSION:: Most of the predictive models produced similar results, and demonstrated the feasibility of identifying individuals with highest probability of having undiagnosed diabetes, through easily-obtained clinical data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Olivera et al. (2017) conducted an observational in Undiagnosed type 2 diabetes. Artificial neural network vs. Logistic regression was evaluated on Mean area under the curve (AUC) in generalization testing. Artificial neural networks and logistic regression models successfully identified individuals with undiagnosed diabetes, achieving mean areas under the curve of 74.17% and 74.41% respectively.

synapsesocial.com/papers/6a1d9bd67328fa9a742fbffehttps://doi.org/10.1590/1516-3180.2016.0309010217
Ask AI
Helpful
Bookmark
Share
View Full Paper