Analysis employs machine learning methods alongside statistical techniques, highlighting predictive performance in employability assessments.
This study investigates how supervised and unsupervised machine learning algorithms can complement traditional statistical methods in the analysis of social survey data. Social science datasets are typically small, noisy, and heterogeneous, which makes robustness and interpretability more important than computational efficiency. Using data from a 2024 survey on the employability of management graduates in Antananarivo, the study compares machine learning approaches with classical multivariate techniques. The objectives are to provide a statistical description of a social reality and to establish criteria for selecting algorithms suited to small-sample contexts. The methodological framework integrates statistical tools such as Chi-square tests, analysis of variance, and multiple regression with exploratory approaches including association rules and clustering. It also incorporates supervised models such as neural networks trained via gradient descent and its variants. Beyond these models, ensemble methods based on decision trees—bagging, random forests, and gradient boosting—are evaluated to highlight their relative strengths. Findings show that gradient boosting offers the most consistent predictive performance while remaining relatively simple to implement. This makes it particularly effective for analysing small and heterogeneous datasets, thereby providing practical value for applied social science research.
No takes yet. Share an insight, caveat, or question.
Hasina et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: