Comparative Evaluation of EfficientNet-B0 and MobileNetV3 for Multi-Class Fire Image Classification with Grad-CAM Interpretation

Comparative Evaluation of EfficientNet-B0 and MobileNetV3 for Multi-Class Fire Image Classification with Grad-CAM Interpretation

Authors

  • Muhammad Fabian Rizky Fatah Department of Informatics Engineering, Universitas Dian Nuswantoro
  • Ravicenna Mahardhika Department of Informatics Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro
  • Muhammad Falah Altairgunna Department of Informatics Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro
  • Ricardus Anggi Pramunendar Department of Informatics Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro

Keywords:

EfficientNet-B0; MobileNetV3; Transfer Learning; Fire Classification; Grad-CAM; Image Classification

Abstract

Fire detection using image-based computer vision is a promising alternative to conventional sensor systems, which are often limited by detection range and false alarms. This study presents a comparative evaluation of two lightweight deep learning architectures, EfficientNet-B0 and MobileNetV3 Small, for multi-class fire image classification across four categories: Urban Fire, Wild Fire, Non-Damage Building, and Non-Damage Wildlife. Both models were trained using transfer learning with ImageNet pre-trained weights on a curated dataset of 1,933 images with selective undersampling to address class imbalance. Grad-CAM (Gradient-weighted Class Activation Mapping) was applied to interpret model decisions and validate the learned visual features. EfficientNet-B0 achieved the highest overall accuracy of 93.8% with a macro F1-score of 0.925, while MobileNetV3 Small reached 91.0% validation accuracy with a macro F1-score of 0.900. Grad-CAM visualizations confirmed that EfficientNet-B0 develops more localized feature activations focused on flame-colored regions, while MobileNetV3 exhibited broader contextual activations. The results demonstrate that lightweight transfer learning models combined with Grad-CAM provide both high classification performance and meaningful interpretability for fire detection applications.

Downloads

Download data is not yet available.

References

M. S. S. Sozol, M. R. H. Mondal, and A. H. Thamrin, “Indoor fire and smoke detection based on optimized YOLOv5,” PLoS One, vol. 20, Apr. 2025, doi: 10.1371/journal.pone.0322052.

M. Faris, E. Ariyanto, and Y. A. S. Yudo, “IMPROVED REAL-TIME HOUSE FIRE DETECTION SYSTEM PERFORMANCE WITH IMAGE CLASSIFICATION USING MOBILENETV2 MODEL,” JIPI (Jurnal Ilm. Penelit. dan Pembelajaran Inform., vol. 8, pp. 656–663, May 2023, doi: 10.29100/jipi.v8i2.3803.

D. L. Nguyen, M. D. Putro, and K. H. Jo, “Lightweight Convolutional Neural Network for Fire Classification in Surveillance System,” IEEE Access, vol. 11, pp. 101604–101615, 2023, doi: 10.1109/ACCESS.2023.3305455.

D. Bhatt et al., “CNN variants for computer vision: History, architecture, application, challenges and future scope,” Oct. 2021, MDPI. doi: 10.3390/electronics10202470.

K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society, Dec. 2016, pp. 770–778. doi: 10.1109/CVPR.2016.90.

C. Szegedy et al., “Going deeper with convolutions,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society, Oct. 2015, pp. 1–9. doi: 10.1109/CVPR.2015.7298594.

C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society, Dec. 2016, pp. 2818–2826. doi: 10.1109/CVPR.2016.308.

G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Institute of Electrical and Electronics Engineers Inc., Nov. 2017, pp. 2261–2269. doi: 10.1109/CVPR.2017.243.

A. Bahtiar, M. I. P. Hutomo, A. Widiyanto, and S. Khomsah, “Class Weighting Approach for Handling Imbalanced Data on Forest Fire Classification Using EfficientNet-B1,” JISKA (Jurnal Inform. Sunan Kalijaga), vol. 10, pp. 63–73, Jan. 2025, doi: 10.14421/jiska.2025.10.1.63-73.

Y. K. Bintang and H. Imaduddin, “PENGEMBANGAN MODEL DEEP LEARNING UNTUK DETEKSI RETINOPATI DIABETIK MENGGUNAKAN METODE TRANSFER LEARNING,” JIPI (Jurnal Ilm. Penelit. dan Pembelajaran Inform., vol. 9, pp. 1442–1455, Aug. 2024, doi: 10.29100/jipi.v9i3.5588.

F. Muhammad, A. B. Elfandra, I. P. Amin, and A. F. Wicaksono, “Pengembangan Model untuk Mendeteksi Kerusakan pada Terumbu Karang dengan Klasifikasi Citra,” Jan. 2026, [Online]. Available: https://arxiv.org/abs/2308.04337

M. Mehmood, A. Shahzad, B. Zafar, A. Shabbir, and N. Ali, “Remote Sensing Image Classification: A Comprehensive Review and Applications,” 2022, Hindawi Limited. doi: 10.1155/2022/5880959.

R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,” Int. J. Comput. Vis., vol. 128, pp. 336–359, Feb. 2020, doi: 10.1007/s11263-019-01228-7.

I. D. Apostolopoulos, I. Athanasoula, M. Tzani, and P. P. Groumpos, “An Explainable Deep Learning Framework for Detecting and Localising Smoke and Fire Incidents: Evaluation of Grad-CAM++ and LIME,” Mach. Learn. Knowl. Extr., vol. 4, pp. 1124–1135, Dec. 2022, doi: 10.3390/make4040057.

F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Institute of Electrical and Electronics Engineers Inc., Nov. 2017, pp. 1800–1807. doi: 10.1109/CVPR.2017.195.

G. M. I. Alam, N. Tasnia, T. Biswas, M. J. Hossen, S. A. Tanim, and M. S. U. Miah, “Real-Time Detection of Forest Fires Using FireNet-CNN and Explainable AI Techniques,” IEEE Access, vol. 13, pp. 51150–51181, 2025, doi: 10.1109/ACCESS.2025.3552352.

H. Yar, F. U. M. Ullah, Z. A. Khan, M. J. Kim, and S. W. Baik, “EFNet-CSM: EfficientNet with a modified attention mechanism for effective fire detection,” Knowledge-Based Syst., vol. 329, Nov. 2025, doi: 10.1016/j.knosys.2025.114353.

J. Simangunsong, M. S. Simanjuntak, and N. D. Simanjuntak, “Analisis Ketepatan Model CNN dalam Deteksi Asap Berbasis Citra,” J. Minfo Polgan, vol. 14, pp. 2381–2390, Dec. 2025, doi: 10.33395/jmp.v14i2.15908.

O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, pp. 211–252, Dec. 2015, doi: 10.1007/s11263-015-0816-y.

C. Shorten and T. M. Khoshgoftaar, “A survey on Image Data Augmentation for Deep Learning,” J. Big Data, vol. 6, Dec. 2019, doi: 10.1186/s40537-019-0197-0.

M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in 36th International Conference on Machine Learning, ICML 2019, International Machine Learning Society (IMLS), 2019, pp. 10691–10700. [Online]. Available: https://proceedings.mlr.press/v97/tan19a.html

A. Howard et al., “Searching for mobileNetV3,” in Proceedings of the IEEE International Conference on Computer Vision, Institute of Electrical and Electronics Engineers Inc., Oct. 2019, pp. 1314–1324. doi: 10.1109/ICCV.2019.00140.

A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Commun. ACM, vol. 60, pp. 84–90, Jun. 2017, doi: 10.1145/3065386.

J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?,” in Advances in Neural Information Processing Systems, Neural information processing systems foundation, 2014, pp. 3320–3328. [Online]. Available: https://papers.nips.cc/paper_files/paper/2014/hash/532a2f85b6977104bc93f8580abbb330-Abstract.html

D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings, International Conference on Learning Representations, ICLR, 2015. [Online]. Available: https://arxiv.org/abs/1412.6980

I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings, International Conference on Learning Representations, ICLR, 2017. [Online]. Available: https://arxiv.org/abs/1608.03983

T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A Next-generation Hyperparameter Optimization Framework,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Association for Computing Machinery, Jul. 2019, pp. 2623–2631. doi: 10.1145/3292500.3330701.

A. Chattopadhyay, A. Sarkar, P. Howlader, and V. N. Balasubramanian, “Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks,” Nov. 2018, doi: 10.1109/WACV.2018.00097.

P. P. Groumpos, “A Critical Historic Overview of Artificial Intelligence: Issues, Challenges, Opportunities, and Threats,” Artif. Intell. Appl., vol. 1, pp. 181–197, Jan. 2023, doi: 10.47852/bonviewAIA3202689.

A. Dosovitskiy et al., “AN IMAGE IS WORTH 16X16 WORDS: TRANSFORMERS FOR IMAGE RECOGNITION AT SCALE,” in ICLR 2021 - 9th International Conference on Learning Representations, International Conference on Learning Representations, ICLR, 2021. [Online]. Available: https://arxiv.org/abs/2010.11929

R. Azad et al., “Advances in medical image analysis with vision Transformers: A comprehensive review,” Jan. 2024, Elsevier B.V. doi: 10.1016/j.media.2023.103000.

Published

2026-06-30

How to Cite

Fatah, M. F. R., Mahardhika, R. ., Altairgunna, M. F. ., & Pramunendar, R. A. . (2026). Comparative Evaluation of EfficientNet-B0 and MobileNetV3 for Multi-Class Fire Image Classification with Grad-CAM Interpretation. Journal of Computing and Smart Ecosystems, 2(1). Retrieved from https://jurnalnew.unimus.ac.id/index.php/J-CaSE/article/view/1129

Issue

Section

Articles
Loading...