Comparative analysis shows competitive accuracy of over 96.8% for stellar classification using machine learning models, highlighting data quality concerns.
Systematic classification of celestial objects is a cornerstone of modern astronomy, as it underpins the understanding of the universe’s large-scale structure and long-term evolution. Moreover, it facilitates the discovery of rare or extreme phenomena, offering clues to the underlying physical laws that govern them. This study presents a comprehensive evaluation of five machine learning models—Random Forest, Adaptive Boosting, Extreme Gradient Boosting, Deep Neural Network, and TabNet—on the Sloan Digital Sky Survey Data Release 18 dataset for the classification of stellar objects into GALAXY, STAR, and Quasi-Stellar Object. With comparative analysis, although the models are not specially tuned, all models achieved competitive performance with accuracy above 96.8%, with Extreme Gradient Boosting reaching an accuracy 98.5%. The reasons for classification are explained by feature importance analysis with SHAP values, which revealed that redshift and the color index E(u−g) are the most influential attributes in distinguishing object classes. Additionally, the study identified several confusing features that contribute to misclassification, suggesting the need for improved data quality in these dimensions. The research highlights the dual role of machine learning in both automation and discovery-driven astronomy.
No takes yet. Share an insight, caveat, or question.
Hanjian Su (2025) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: