Accuracy-Efficiency Benchmarking of Lightweight Machine Learning Models for Building Heating and Cooling Load Prediction
Keywords:
Building Energy Efficiency, Heating Load, Cooling Load, Lightweight Machine Learning, Gradient Boosting, Smart BuildingAbstract
Early-stage estimation of heating and cooling loads supports energy-efficient building design, but complex predictive models may impose unnecessary computational costs for small tabular datasets. This study benchmarks four lightweight regression models, such as Linear Regression, Ridge Regression, Random Forest, and Gradient Boosting, using the UCI Energy Efficiency dataset containing 768 simulated building configurations, eight design variables, and two continuous targets. Model accuracy was evaluated with shuffled 10-fold cross-validation using mean absolute error (MAE), root mean squared error (RMSE), and the coefficient of determination (R²). Computational efficiency was assessed through training time, prediction latency, and serialized model size, while permutation importance was used for interpretation. Random Forest achieved the lowest heating-load RMSE (0.4624) and an R² of 0.9978; however, its difference from Gradient Boosting was not statistically significant. Gradient Boosting was approximately 96 times smaller and 20 times faster at inference. For cooling load, Gradient Boosting achieved the lowest RMSE (1.4903) and an R² of 0.9751, significantly outperforming Random Forest in fold-level RMSE. Relative compactness was the most influential heating-load feature, whereas overall height dominated cooling-load prediction. The results indicate that Gradient Boosting offers the strongest overall accuracy-efficiency trade-off for lightweight smart-building prediction systems.
Downloads
References
United Nations Environment Programme and Global Alliance for Buildings and Construction, Global Status Report for Buildings and Construction 2024/2025. Nairobi, Kenya: UNEP, 2025. [Online]. Available: https://www.unep.org/resources/report/global-status-report-buildings-and-construction-20242025
International Energy Agency, “Buildings,” in Energy Efficiency 2025. Paris, France: IEA, 2025. [Online]. Available: https://www.iea.org/reports/energy-efficiency-2025/buildings
C. Fan, D. Yan, F. Xiao, A. Li, J. An, and X. Kang, “Advanced data analytics for enhancing building performances: From data-driven to big data-driven approaches,” Building Simulation, vol. 14, no. 1, pp. 3–24, 2021, doi: 10.1007/s12273-020-0723-1.
L. Wederhake, S. Wenninger, C. Wiethe, and G. Fridgen, “On the surplus accuracy of data-driven energy quantification methods in the residential sector,” Energy Informatics, vol. 5, art. no. 7, 2022, doi: 10.1186/s42162-022-00194-8.
A. Tsanas and A. Xifara, “Accurate quantitative estimation of energy performance of residential buildings using statistical machine learning tools,” Energy and Buildings, vol. 49, pp. 560–567, 2012, doi: 10.1016/j.enbuild.2012.03.003.
A. Tsanas and A. Xifara, “Energy Efficiency,” UCI Machine Learning Repository, 2012, doi: 10.24432/C51307.
R. Chaganti, F. Rustam, T. Daghriri, I. de la Torre Díez, J. L. V. Mazón, C. L. Rodríguez, and I. Ashraf, “Building heating and cooling load prediction using ensemble machine learning model,” Sensors, vol. 22, no. 19, art. no. 7692, 2022, doi: 10.3390/s22197692.
J. Guo, S. Yun, Y. Meng, N. He, D. Ye, Z. Zhao, et al., “Prediction of heating and cooling loads based on light gradient boosting machine algorithms,” Building and Environment, vol. 236, art. no. 110252, 2023, doi: 10.1016/j.buildenv.2023.110252.
G. Bekdaş, Y. Aydın, Ü. Işıkdağ, A. N. Sadeghifam, S. Kim, and Z. W. Geem, “Prediction of cooling load of tropical buildings with machine learning,” Sustainability, vol. 15, no. 11, art. no. 9061, 2023, doi: 10.3390/su15119061.
Y. Chen, Y. Ye, J. Liu, L. Zhang, W. Li, and S. Mohtaram, “Machine learning approach to predict building thermal load considering feature variable dimensions: An office building case study,” Buildings, vol. 13, no. 2, art. no. 312, 2023, doi: 10.3390/buildings13020312.
Y. Zhou, Y. Liu, D. Wang, and X. Liu, “Comparison of machine-learning models for predicting short-term building heating load using operational parameters,” Energy and Buildings, vol. 253, art. no. 111505, 2021, doi: 10.1016/j.enbuild.2021.111505.
M. Rana, S. Sethuvenkatraman, and M. Goldsworthy, “A data-driven approach based on quantile regression forest to forecast cooling load for commercial buildings,” Sustainable Cities and Society, vol. 76, art. no. 103511, 2022, doi: 10.1016/j.scs.2021.103511.
W. Gao, X. Huang, M. Lin, J. Jia, and Z. Tian, “Short-term cooling load prediction for office buildings based on feature selection scheme and stacking ensemble model,” Engineering Computations, vol. 39, no. 5, pp. 2003–2029, 2022, doi: 10.1108/EC-07-2021-0406.
Y. Shen, Y. Hu, K. Cheng, H. Yan, K. Cai, J. Hua, et al., “Utilizing interpretable stacking ensemble learning and NSGA-III for the prediction and optimisation of building photo-thermal environment and energy consumption,” Building Simulation, vol. 17, pp. 819-838, 2024, doi: 10.1007/s12273-024-1108-7.
B. Chegari, M. Tabaa, E. Simeu, F. Moutaouakkil, and H. Medromi, “Multi-objective optimization of building energy performance and indoor thermal comfort by combining artificial neural networks and metaheuristic algorithms,” Energy and Buildings, vol. 239, art. no. 110839, 2021, doi: 10.1016/j.enbuild.2021.110839.
C. Fan, F. Xiao, C. Yan, C. Liu, Z. Li, and J. Wang, “A novel methodology to explain and evaluate data-driven building energy performance models based on interpretable machine learning,” Applied Energy, vol. 235, pp. 1551-1560, 2019, doi:10.1016/j.apenergy.2018.11.081.
L. Breiman, “Random forests,” Machine Learning, vol. 45, pp. 5–32, 2001, doi:10.1023/A:1010933404324.
J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001, doi: 10.1214/aos/1013203451.
A. E. Hoerl and R. W. Kennard, “Ridge regression: Biased estimation for nonorthogonal problems,” Technometrics, vol. 12, no. 1, pp. 55–67, 1970, doi:10.1080/00401706.1970.10488634.
F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” Journal of Machine Learning Research, vol. 7, pp. 1–30, 2006.
S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65–70, 1979.
A. Altmann, L. Toloşi, O. Sander, and T. Lengauer, “Permutation importance: A corrected feature importance measure,” Bioinformatics, vol. 26, no. 10, pp. 1340–1347, 2010, doi: 10.1093/bioinformatics/btq134.
F. Wilcoxon, “Individual comparisons by ranking methods,” Biometrics Bulletin, vol. 1, no. 6, pp. 80–83, 1945, doi: 10.2307/3001968.
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Journal of Computing and Smart Ecosystems

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.