PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026Scientific Reports2 citationsOpen Access

Enhancing logistic regression classification: insights from simulation and real-world applications through ranked set sampling

RYRazieh YousefiBLBenoit LiquetMMMahdi Mahdizadeh

Key Points

  • This research aims to evaluate the effectiveness of ranked set sampling methods in enhancing logistic regression classification performance in healthcare contexts.
  • Conducted simulation studies comparing performance of logistic regression under ranked set sampling (RSS), extreme ranked set sampling (ERSS), and simple random sampling (SRS).
  • Tested approaches on datasets predicting osteoporosis and maternal health risks.
  • Assessed models based on various sample sizes, population sizes, and class imbalance ratios.
  • ERSS generally outperformed SRS and RSS, especially in small or imbalanced datasets.
  • At set sizes k of 5 and 10, ERSS showed significantly better performance metrics compared to the alternatives, achieving scores above 0.95.
  • Performance improved with increasing sample size, with ERSS consistently demonstrating the best results.

Abstract

Machine learning algorithms are widely used for disease prediction, but the choice of sampling method is a factor that affects model performance, particularly in large populations, limitations on the availability of resources, and imbalanced datasets, which is common in healthcare research. Ranked set sampling (RSS) offers a more efficient alternative to simple random sampling (SRS) when precise measurement requires great effort and determination, yet ranking sampling units is achievable. In this study, logistic regression (LR) is employed as the most commonly used model in machine learning, and its performance is evaluated under RSS and extreme ranked set sampling (ERSS). We conducted a comprehensive comparison based on a simulation study to evaluate the performance of LR under RSS, ERSS, and SRS across different population sizes, sample sizes, and class imbalance ratios. Set size and correlation between the response variable and the concomitant variable, which is used for ranking, were two key sampling design characteristics examined in RSS and ERSS. The models’ effectiveness was assessed through several standard metrics. Furthermore, the approaches were tested on two practical datasets: one focused on predicting osteoporosis, and the other on assessing maternal health risks. ERSS generally performed better than SRS and RSS, particularly in small or imbalanced datasets, with the advantage becoming more pronounced as the set size (k) increased. When CDATA[k = 3], the methods performed similarly. When CDATA[k = 5], ERSS often outperformed the other methods, and when CDATA[k = 10], ERSS performed best in most scenarios and its metrics achieved higher than 0.95 in most scenarios. With an increase in sample size, the performance of all models improved, but ERSS still exhibited the best performance. On the other hand, ERSS also demonstrated decent performance at small sample sizes based on most metrics. The change in population size did not have a substantial impact on the results. Despite this good performance of ERSS, RSS often exhibited similar performance compared to SRS. In the osteoporosis dataset, ERSS had better LR performance when CDATA[k=5] or 10. In the maternal health risk dataset, ERSS improved LR performance, especially when sampling was based on a concomitant variable with higher correlation, validating the simulation findings and demonstrating the practical benefits of rank-based methods. The findings indicated that RSS provided comparable performance to SRS in certain conditions. But, ERSS was highly beneficial compared to SRS and RSS in terms of achieving higher performance metrics. It achieved better performance in both balanced and imbalanced situations, and also with limited sample sizes. Its robustness and efficiency make it a valuable tool in different studies. However, these findings should be used considering the study’s limitations, such as constraints in evaluating extreme imbalance scenarios (e.g., when the correlation coefficient is less than 0.1) and the evaluation of just RSS and ERSS designs. Real-world datasets have also confirmed ERSS’s ability to improve LR’s performance, consistent with simulation results. Overall, our findings indicated that rank-based sampling methods produce results that are at least as favorable as those of SRS, and, given suitable settings, can improve classification performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yousefi et al. (2026) studied this question.

synapsesocial.com/papers/69aa6f3c531e4c4a9ff5955dhttps://doi.org/10.1038/s41598-026-41333-5
Ask AI
Helpful
Bookmark
Share
View Full Paper