Multi-stage framework combines dimensionality reduction and ensemble techniques, achieving top performance in a competitive setting.
We present a multi-stage framework for predictive modeling that integrates automated feature engineering, selective dimensionality reduction, and targeted ensembling. Our pipeline begins with feature generation using a GPU-accelerated adaptation of AutoFeat, followed by variance-based pruning and LightGBM gain-based selection to isolate informative predictors for each target prop- erty. Target-specific feature subsets were then optimized through supervised cutoff sweeps, evaluated with TabPFNRegressor. To create a diverse ensemble, we trained three distinct TabPFN pipelines on feature sets of progressive refinement: the full engineered space, a variance-pruned subset, and the final target-specific optimized features. Final predictions were produced via property-level en- sembling, blending outputs from multiple TabPFN models. This approach achieved a final rank of 8th on the public leaderboard of the Shell.ai Hackathon 2025 with a score of 95.56 [1] on the test set, placing it in the top percentile of over 1,000 teams. The resulting models were deployed in an interactive web-based prototype to demonstrate their practical utility.
No takes yet. Share an insight, caveat, or question.
Daramola et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: