PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 25, 2026PLOS Digital Health0 citationsOpen Access

Classification of knowledge of fertility period among adolescent girls in East Africa from 2012 to 2022: Machine learning algorithm

View Full Paper
ABAndualem Addisu BirlieKGKassahun Dessie GashuMKMulugeta Desalegn Kasaye

Key Points

  • The study aims to classify knowledge of the fertility period among adolescent girls in East Africa from 2012 to 2022 using machine learning algorithms.
  • Community-based cross-sectional study design utilizing DHS datasets from 12 East African countries.
  • Application of ten machine learning algorithms for classification and identification of predictors.
  • Data analysis performed using R and Python with techniques such as data cleaning, one-hot encoding, and ten-fold cross-validation.
  • Only 13.22% of the 40,664 adolescent girls were knowledgeable about their fertility period.
  • Logistic regression achieved 74.38% AUC and 82.71% accuracy on unbalanced training data.
  • Random forest outperformed with 91.12% AUC and 83.26% accuracy on balanced training data.

Abstract

Understanding the time of the menstrual cycle would help women to avoid getting pregnant without the need for surgical, hormonal, or mechanical contraception. Women who do not use contraception and do not know when they are fertile are at a higher risk (17%) of unplanned pregnancy and abortion. Classifying knowledge of fertility periods using machine learning algorithms would help to automate decision-making, produce more precise and accurate classification, and scale up to manage big and complex datasets. Therefore, this study aimed to classify knowledge of the fertility period among adolescent girls in East Africa from 2012 to 2022 using a machine-learning algorithm. A community-based cross-sectional study design was used from 12 East African countries’ DHS datasets spanning 2012–2022. The machine learning algorithms were applied to classify knowledge of the fertility period and identify its predictors using R software and Python, particularly Jupiter Notebook in Anaconda. Data cleaning, one-hot encoding, data splitting, data balancing, and ten-fold cross-validation were performed. Ten machine learning algorithms and SHAP were used to select and interpret the best model. From the 40,664 adolescent girls in East Africa, 13.22% (95% CI: 12.91, 13.54) of participants had knowledge of the fertility period. Logistic regression was found to be the best model for unbalanced training data with 74.38% of an AUC and 82.71% of an accuracy. While random forest outperformed on balanced training data, it achieved 91.12% of an AUC and 83.26% accuracy. The key determinant factors of the knowledge of the fertility period were education level, country, hearing about family planning, hearing about sexually transmitted infections, wealth index, knowledge of any method, and visiting health facilities. Governments, NGOs, policy makers, and researchers can utilize these findings to design targeted interventions for improving adolescents’ reproductive health based on the identified gaps and disparities.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Birlie et al. (2026) studied this question.

synapsesocial.com/papers/699e911bf5123be5ed04e770https://doi.org/10.1371/journal.pdig.0001108
Ask AI
Helpful
Bookmark
Share
View Full Paper