Machine Learning Prediction of Concrete Compressive Strength: Model Comparison, CatBoost Optimization, and SHAP Interpretation

Authors

  • Musthafa 'Abduh Fakhruddin Teknik Informatika, Fakultas Ilmu Komputer, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Sri Winarno Teknik Informatika, Fakultas Ilmu Komputer, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Acun Kardianawati Sistem Informasi, Fakultas Ilmu Komputer, Universitas Dian Nuswantoro, Semarang, Indonesia

DOI:

https://doi.org/10.36080/idealis.v9i2.3848

Keywords:

Bayesian optimization, CatBoost, Concrete compressive strength, SHAP interpretability, Machine Learning

Abstract

Accurate prediction of concrete compressive strength is vital for structural design, yet conventional testing is constrained by lengthy curing requirements. Machine learning offers an alternative by modeling non-linear mix-performance interactions. This study presents a comparative framework evaluating nine regression algorithms using the UCI Concrete Compressive Strength dataset (1,005 samples). Performance was assessed via 10x5 repeated cross-validation with 95% confidence intervals, and statistical significance was evaluated using a Linear Mixed-Effects Model with Holm-Bonferroni corrected pairwise t-tests. Tree-based ensembles outperformed linear approaches, with CatBoost yielding the highest baseline cross-validation R² of 0.931 (95% CI: 0.927 to 0.935). Subsequent Bayesian hyperparameter optimization via Optuna’s Tree-structured Parzen Estimator (400 trials) improved the final CatBoost model’s performance to a test  of 0.943, RMSE of 4.142 MPa, and MAE of 2.616 MPa. SHAP analysis indicated that curing age, the water-to-binder ratio, and cement are the dominant predictors, while the model exhibited physically consistent behavior aligned with concrete hydration kinetics and Abrams' law. This work's key contribution is jointly integrating correlation-corrected statistical validation, multi-model Bayesian optimization, and domain-informed feature engineering with SHAP interpretation, rarely combined in prior concrete-strength studies. The framework offers an accurate, interpretable tool for preliminary concrete mix design.

Downloads

Download data is not yet available.

References

[1] M. Shaaban, M. Amin, S. Selim, and I. M. Riad, “Machine learning approaches for forecasting compressive strength of high-strength concrete,” Sci. Rep., vol. 15, no. 1, p. 25567, 2025, doi: 10.1038/s41598-025-10342-1.

[2] J. Huang, M. M. S. Sabri, D. V. Ulrikh, M. Ahmad, and K. A. M. Alsaffar, “Predicting the Compressive Strength of the Cement-Fly Ash–Slag Ternary Concrete Using the Firefly Algorithm (FA) and Random Forest (RF) Hybrid Machine-Learning Method,” Materials, vol. 15, no. 12, p. 4193, 2022, doi: 10.3390/ma15124193.

[3] H. Fu, X. Zhou, P. Xu, and D. Sun, “Prediction of Compressive Strength of Concrete Using Explainable Machine Learning Models,” Materials, vol. 18, no. 21, p. 5009, 2025, doi: 10.3390/ma18215009.

[4] A. K. Sah and Y.-M. Hong, “Performance Comparison of Machine Learning Models for Concrete Compressive Strength Prediction,” Materials, vol. 17, no. 9, p. 2075, 2024, doi: 10.3390/ma17092075.

[5] I.-C. Yeh, “Modeling of strength of high-performance concrete using artificial neural networks,” Cem. Concr. Res., vol. 28, no. 12, pp. 1797–1808, 1998, doi: 10.1016/S0008-8846(98)00165-3.

[6] I. N. Fathy, H. A. Dahish, M. K. Alkharisi, A. A. Mahmoud, and H. E. E. Fouad, “Predicting the compressive strength of concrete incorporating waste powders exposed to elevated temperatures utilizing machine learning,” Sci. Rep., vol. 15, no. 1, p. 25275, 2025, doi: 10.1038/s41598-025-11239-9.

[7] X. Chen, X. Zhang, and W.-Z. Chen, “Advanced Predictive Modeling of Concrete Compressive Strength and Slump Characteristics: A Comparative Evaluation of BPNN, SVM, and RF Models Optimized via PSO,” Materials, vol. 17, no. 19, p. 4791, 2024, doi: 10.3390/ma17194791.

[8] M. F. Javed, M. Fawad, R. Lodhi, T. Najeh, and Y. Gamil, “Forecasting the strength of preplaced aggregate concrete using interpretable machine learning approaches,” Sci. Rep., vol. 14, no. 1, p. 8381, 2024, doi: 10.1038/s41598-024-57896-0.

[9] S. Elhishi, A. M. Elashry, and S. El-Metwally, “Unboxing machine learning models for concrete strength prediction using XAI,” Sci. Rep., vol. 13, no. 1, p. 19892, 2023, doi: 10.1038/s41598-023-47169-7.

[10] O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Sci. Rep., vol. 14, no. 1, p. 6086, 2024, doi: 10.1038/s41598-024-56706-x.

[11] J. Liu and Y. Xu, “T-Friedman Test: A New Statistical Test for Multiple Comparison with an Adjustable Conservativeness Measure,” International Journal of Computational Intelligence Systems, vol. 15, no. 1, p. 29, 2022, doi: 10.1007/s44196-022-00083-8.

[12] J. M. Gorriz, J. Ramirez, F. Segovia, C. Jimenez-Mesa, F. J. Martinez-Murcia, and J. Suckling, “Statistical agnostic regression: A machine learning method to validate regression models,” J. Adv. Res., vol. 80, pp. 503–533, 2026, doi: 10.1016/j.jare.2025.04.026.

[13] L. W. Rizkallah, “Enhancing the performance of gradient boosting trees on regression problems,” J. Big Data, vol. 12, no. 1, p. 35, 2025, doi: 10.1186/s40537-025-01071-3.

[14] J. Ma et al., “A comprehensive comparison among metaheuristics (MHs) for geohazard modeling using machine learning: Insights from a case study of landslide displacement prediction,” Eng. Appl. Artif. Intell., vol. 114, p. 105150, 2022, doi: 10.1016/j.engappai.2022.105150.

[15] R. S. Ajin, S. Segoni, and R. Fanti, “Optimization of SVR and CatBoost models using metaheuristic algorithms to assess landslide susceptibility,” Sci. Rep., vol. 14, no. 1, p. 24851, 2024, doi: 10.1038/s41598-024-72663-x.

[16] P.-O. Côté, A. Nikanjam, N. Ahmed, D. Humeniuk, and F. Khomh, “Data cleaning and machine learning: a systematic literature review,” Automated Software Engineering, vol. 31, no. 2, p. 54, 2024, doi: 10.1007/s10515-024-00453-w.

[17] F. C. Oettl, J. F. Oeding, R. Feldt, C. Ley, M. T. Hirschmann, and K. Samuelsson, “The artificial intelligence advantage: Supercharging exploratory data analysis,” Knee Surgery, Sports Traumatology, Arthroscopy, vol. 32, no. 11, pp. 3039–3042, 2024, doi: 10.1002/ksa.12389.

[18] Y. Zhang, W. Ren, Y. Chen, Y. Mi, J. Lei, and L. Sun, “Predicting the compressive strength of high-performance concrete using an interpretable machine learning model,” Sci. Rep., vol. 14, no. 1, p. 28346, 2024, doi: 10.1038/s41598-024-79502-z.

[19] A. M. Neville, Properties of Concrete, 5th ed. Harlow, UK: Prentice Hall, 2012.

[20] D. Sofia and P. Sekarpuji, “Penerapan Metode Stacking Ensemble Untuk Analisis Sentimen Pada Ulasan Aplikasi Ruang Guru,” IDEALIS: InDonEsiA journaL Information System, vol. 8, no. 2, pp. 248–257, 2025, doi: 10.36080/idealis.v8i2.3559.

[21] P. Li, Z. Chen, X. Chu, and K. Rong, “DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data,” Proceedings of the ACM on Management of Data, vol. 1, no. 2, pp. 1–26, 2023, doi: 10.1145/3589328.

[22] M. A. Bouke and A. Abdullah, “An empirical study of pattern leakage impact during data preprocessing on machine learning-based intrusion detection models reliability,” Expert Syst. Appl., vol. 230, p. 120715, 2023, doi: 10.1016/j.eswa.2023.120715.

[23] P. Probst, A.-L. Boulesteix, and B. Bischl, “Tunability: importance of hyperparameters of machine learning algorithms,” Journal of Machine Learning Research, vol. 20, no. 53, pp. 1–32, 2019, Accessed: Jul. 14, 2026. [Online]. Available: https://jmlr.org/papers/volume20/18-444/18-444.pdf

[24] F. Shehzad, T. Breuer, M. Maistro, and D. Jannach, “‘We Share Our Code Online’: Why This Is Not Enough to Ensure Reproducibility and Progress in Recommender Systems Research,” in Proceedings of the Nineteenth ACM Conference on Recommender Systems, Prague, Czech Republic: Association for Computing Machinery, 2025, pp. 884–893. doi: 10.1145/3705328.3748157.

[25] H. Ahmed and J. Lofstead, “Managing Randomness to Enable Reproducible Machine Learning,” in Proceedings of the 5th International Workshop on Practical Reproducible Evaluation of Computer Systems, Minneapolis, MN, USA, 2022, pp. 15–20. doi: 10.1145/3526062.3536353.

[26] H. Semmelrock et al., “Reproducibility in machine‐learning‐based research: Overview, barriers, and drivers,” AI Mag., vol. 46, no. 2, p. e70002, 2025, doi: 10.1002/aaai.70002.

[27] D. Chicco, M. J. Warrens, and G. Jurman, “The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation,” PeerJ Comput. Sci., vol. 7, p. e623, 2021, doi: 10.7717/peerj-cs.623.

[28] W. G. Mokodaser, H. Koapaha, and S. I. Adam, “Model Random Forest Data Historis Multivariat Untuk Prediksi Pendapatan Asuransi,” IDEALIS: InDonEsiA journaL Information System, vol. 8, no. 2, pp. 267–276, 2025, doi: 10.36080/idealis.v8i2.3512.

[29] Y. M. Jaradat, M. A. Alia, M. Z. Masoud, A. A. Manasrah, I. A. Jannoud, and O. Alheyasat, “Beyond One-Size-Fits-All: Comparing and Selecting Regression Metrics for Robust Model Assessment,” in 2025 12th International Conference on Information Technology (ICIT), Amman, Jordan, 2025, pp. 416–422. doi: 10.1109/ICIT64950.2025.11049268.

[30] I. Hussain, K. B. Ching, C. Uttraphan, K. G. Tay, and A. Noor, “Evaluating machine learning algorithms for energy consumption prediction in electric vehicles: A comparative study,” Sci. Rep., vol. 15, no. 1, p. 16124, 2025, doi: 10.1038/s41598-025-94946-7.

[31] W. Li, D. Cook, E. Tanaka, and S. VanderPlas, “A Plot is Worth a Thousand Tests: Assessing Residual Diagnostics with the Lineup Protocol,” Journal of Computational and Graphical Statistics, vol. 33, no. 4, pp. 1497–1511, 2024, doi: 10.1080/10618600.2024.2344612.

[32] R. T. Nakatsu, “Validation of machine learning ridge regression models using Monte Carlo, bootstrap, and variations in cross-validation,” Journal of Intelligent Systems, vol. 32, no. 1, p. 20220224, 2023, doi: 10.1515/jisys-2022-0224.

[33] S. Bates, T. Hastie, and R. Tibshirani, “Cross-Validation: What Does It Estimate and How Well Does It Do It?,” J. Am. Stat. Assoc., vol. 119, no. 546, pp. 1434–1445, 2024, doi: 10.1080/01621459.2023.2197686.

[34] V. A. Brown, “An Introduction to Linear Mixed-Effects Modeling in R,” Adv. Methods Pract. Psychol. Sci., vol. 4, no. 1, p. 2515245920960351, 2021, doi: 10.1177/2515245920960351.

[35] Z. Yu, M. Guindani, S. F. Grieco, L. Chen, T. C. Holmes, and X. Xu, “Beyond t test and ANOVA: applications of mixed-effects models for more rigorous statistical analysis in neuroscience research,” Neuron, vol. 110, no. 1, pp. 21–35, 2022, doi: 10.1016/j.neuron.2021.10.030.

[36] C. Nadeau and Y. Bengio, “Inference for the Generalization Error,” Mach. Learn., vol. 52, no. 3, pp. 239–281, 2003, doi: 10.1023/A:1024068626366.

[37] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Comput., vol. 10, no. 7, pp. 1895–1923, 1998, doi: 10.1162/089976698300017197.

[38] R. R. Bouckaert and E. Frank, “Evaluating the Replicability of Significance Tests for Comparing Learning Algorithms,” in Advances in Knowledge Discovery and Data Mining, Sydney, Australia, 2004, pp. 3–12. doi: https://doi.org/10.1007/978-3-540-24775-3_3.

[39] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A Next-generation Hyperparameter Optimization Framework,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 2019, pp. 2623–2631. doi: 10.1145/3292500.3330701.

[40] S. Ali, F. Akhlaq, A. S. Imran, Z. Kastrati, S. M. Daudpota, and M. Moosa, “The enlightening role of explainable artificial intelligence in medical & healthcare domains: A systematic literature review,” Comput. Biol. Med., vol. 166, p. 107555, 2023, doi: 10.1016/j.compbiomed.2023.107555.

[41] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, in NIPS’17. Long Beach, California, USA, 2017, pp. 4768–4777.

[42] S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees,” Nat. Mach. Intell., vol. 2, no. 1, pp. 56–67, 2020, doi: 10.1038/s42256-019-0138-9.

[43] Z. Wang, H. Liu, M. N. Amin, K. Khan, M. T. Qadir, and S. A. Khan, “Optimizing machine learning techniques and SHapley Additive exPlanations (SHAP) analysis for the compressive property of self-compacting concrete,” Mater. Today Commun., vol. 39, p. 108804, 2024, doi: 10.1016/j.mtcomm.2024.108804.

[44] P. Novello, G. Poëtte, D. Lugato, and P. M. Congedo, “Goal-Oriented Sensitivity Analysis of Hyperparameters in Deep Learning,” J. Sci. Comput., vol. 94, no. 3, p. 45, 2023, doi: 10.1007/s10915-022-02083-4.

[45] Y. T. Altuncı, “A Comprehensive Study on the Estimation of Concrete Compressive Strength Using Machine Learning Models,” Buildings, vol. 14, no. 12, p. 3851, 2024, doi: 10.3390/buildings14123851.

Downloads

Published

07/31/2026

How to Cite

[1]
M. ’Abduh Fakhruddin, S. Winarno, and A. Kardianawati, “Machine Learning Prediction of Concrete Compressive Strength: Model Comparison, CatBoost Optimization, and SHAP Interpretation”, IDEALIS, vol. 9, no. 2, pp. 393–406, Jul. 2026.