The increasing complexity of scientific research projects and growing competition for public funding require data-driven approaches to support evidence-based decision-making in research management. This paper presents an integrated machine learning-based analytical pipeline for scientific project analysis, positioning the contribution as methodological integration and empirical insight rather than the development of new algorithms. Using open data from the CORDIS database on projects funded under the Sixth Framework Programme (FP6), the study combines three complementary analytical tasks—structure identification, budget prediction, and funding-completeness classification—within a single empirical decision-support workflow. Unsupervised analysis based on dimensionality reduction and clustering reveals distinct structural patterns in project characteristics and identifies atypical large-scale projects characterized by substantially higher budgets and consortium sizes. Regression models predict total project costs, achieving R2 ≈ 0.50 on unseen data; the gap between explained and unexplained variance is interpreted in relation to latent contextual factors not available in administrative records. Classification models distinguish between fully and partially funded projects (threshold r ≥ 0.70); the high discriminative performance (AUC > 0.97) is interpreted cautiously in light of target-variable construction and potential feature dependence. Association rule mining identifies interpretable funding patterns (strongest rule lift = 6.30). The study contributes a reproducible example of how mature machine learning techniques can be systematically integrated to generate domain-specific evidence for research funding governance.
Medetbek et al. (Tue,) studied this question.