PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026Buildings2 citationsOpen Access

Exploring the Impact of Different Clustering Algorithms on the Performance of Ensemble Learning-Based Mass Appraisal Models

View Full Paper
SŞSüleyman ŞişmanAKAbdullah KaraAAArif Çağdaş Aydınoğlu

Key Points

  • The research aims to evaluate the effects of different clustering algorithms on ensemble learning models for mass appraisal accuracy.
  • Evaluated clustering algorithms including K-Means, K-Medians, and SCMCA.
  • Used a comprehensive real estate dataset for analysis.
  • Assessed clustering quality using Silhouette, Calinski–Harabasz, and Davies–Bouldin indices.
  • Compared performance of ensemble models like Random Forest, GBM, XGBoost, and LightGBM with clustering.
  • SCMCA combined with LightGBM achieved the best performance with RMSE = 0.061 and R2 = 0.722.
  • Clustering improved MAE by up to 7.26%, MAPE by 10.61%, and RMSE by 8.40%.
  • Clustering effectively enhances predictive performance and quality of mass appraisal models.

Abstract

Mass appraisal models are gaining use for improving valuation accuracy, yet their performance remains highly sensitive to how spatial and non-spatial data are structured before training. Clustering algorithms can be used to segment heterogeneous property groups into more homogeneous ones, potentially improving predictive performance. This study investigates the impact of different clustering algorithms, (i.e., K-Means, K-Medians and the Spatially Constrained Multivariate Clustering Algorithm (SCMCA)), on the performance of prominent ensemble learning-based mass appraisal models (i.e., Random Forest (RF), the Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost) and the Light Gradient Boosting Machine (LightGBM)). Using a comprehensive real estate dataset, clustering quality is evaluated using Silhouette, Calinski–Harabasz, and Davies–Bouldin indices, and the performance of cluster-based ensemble mass appraisal models is then compared. The findings indicate that the best performance is achieved with the SCMCA–LightGBM model combination, which reached RMSE = 0.061 and R2 = 0.722. Furthermore, it is determined that clustering-based models provide improvements of up to 7.26% in MAE, 10.61% in MAPE, and 8.40% in RMSE, depending on the combination. The results show that clustering is an effective preprocessing step that can substantially enhance the predictive performance and overall quality of mass appraisal models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Şişman et al. (2026) studied this question.

synapsesocial.com/papers/698434ebf1d9ada3c1fb39ebhttps://doi.org/10.3390/buildings16030615
Ask AI
Helpful
Bookmark
Share
View Full Paper