PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 2026Journal of Composites Science0 citationsOpen Access

Data-Driven Prediction of Compressive Strength in Concrete with Lightweight Expanded Clay Aggregate Using Machine Learning Techniques

SNSoorya M. NairKarunya UniversityANAnand NammalvarKarunya UniversityAAA. Diana AndrushiaKarunya University

Key Points

  • The paper aims to develop machine learning models to predict the compressive strength of LECA concrete accurately.
  • Created and preprocessed a dataset using statistical normalization and correlation analysis.
  • Developed five supervised machine learning models: MLR, SVR, RF, XGBoost, and CatBoost.
  • Fine-tuned models using grid-search strategy and ten-fold cross-validation.
  • Evaluated prediction quality using metrics like R2, RMSE, MAE, and MAPE.
  • Compared models using the Gray Relational Analysis (GRA) method.
  • CatBoost outperformed other models with an R2 of 0.907, RMSE of 3.41 MPa, MAE of 2.47 MPa, and MAPE of 10.05%.
  • Identified water and LECA content as the most significant factors influencing compressive strength.
  • The gradient boosting model effectively captured nonlinear interactions in LECA concrete.

Abstract

The growing need for sustainable and lightweight building materials has accelerated research on alternatives to conventional concretes, out of which Lightweight Expanded Clay Aggregate (LECA) concrete has emerged as a promising solution. However, the high porosity and nonlinear mechanical behavior of LECA concrete complicate the accurate prediction of compressive strength through conventional empirical models. The main focus of the paper is on identifying a comprehensive machine learning-based framework for modeling and predicting the 28-day compressive strength of LECA-based lightweight concrete. The dataset was created and preprocessed by using statistical normalization and correlation analysis. In this study, five supervised machine learning models—Multiple Linear Regression (MLR), Support Vector Regression (SVR), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Categorical Boosting (CatBoost)—were developed and fine-tuned using a grid-search strategy combined with ten-fold cross-validation. The quality of the prediction made by each model was evaluated by means of standard performance indicators, such as the coefficient of determination (R2), root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE). After the evaluation, the models were subsequently compared and ranked according to the Gray Relational Analysis (GRA) method. The comparative assessment shows that CatBoost demonstrated the most reliable performance, achieving an R2 of 0.907, RMSE of 3.41 MPa, MAE of 2.47 MPa, and MAPE of 10.05%, outperforming the remaining algorithms. To interpret the significance of features, SHAP (Shapley Additive exPlanations) analysis was applied, which identified water and LECA content as the dominant factors influencing compressive strength, followed by the cement and fine aggregate proportions. The findings reveal that the ensemble-based gradient boosting model is capable of capturing intricate nonlinear interactions, as observed in the heterogeneous matrix of LECA concrete.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nair et al. (2026) studied this question.

synapsesocial.com/papers/69b258a396eeacc4fcec88d6https://doi.org/10.3390/jcs10030151
Ask AI
Helpful
Bookmark
Share
View Full Paper