Improving the energy performance of residential buildings requires models that are not only accurate but also transparent and interpretable. This study applies multiple state-of-the-art Symbolic Regression (SR) algorithms to the well-known Energy Efficiency Dataset to predict Heating Load (HL) and Cooling Load (CL) based on architectural features. Unlike previous studies that rely heavily on black-box machine learning methods, this work generates explicit, human-readable analytical expressions that reveal the underlying relationships between building parameters and thermal demands. The proposed SR models achieve competitive accuracy, with RMSE values of 1.16 (HL) and 2.35 (CL), and R² values of 0.987 and 0.941, respectively. In addition to strong quantitative performance, the symbolic expressions provide qualitative insights into feature importance, such as the dominant role of compactness, surface area, and glazing characteristics. The novelty of this study lies in offering the first comprehensive symbolic regression framework for this dataset, presenting interpretable analytical equations that preserve physical meaning, and comparing SR directly with Linear Regression, Random Forest, and Gradient Boosting to demonstrate that interpretability can be achieved without compromising predictive accuracy. These results highlight the potential of symbolic modelling as a transparent and reliable tool for energy-efficient building design.
Öztürk et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: