Comparative Analysis of Detection and Segmentation Models for Building Damage Assessment from UAV Imagery
DOI:
https://doi.org/10.31861/sisiot2026.1.01012Keywords:
object detection, semantic segmentation, damage assessment, UAV, deep learningAbstract
This paper presents a systematic comparison of object detection models (Faster R-CNN, RetinaNet, YOLOv8) and semantic segmentation models (U-Net, DeepLabV3+, SegFormer) for automated building damage assessment from unmanned aerial vehicle (UAV) imagery. The comparison covers both two-stage (Faster R-CNN) and single-stage (RetinaNet, YOLOv8) approaches to detection, as well as convolutional (U-Net, DeepLabV3+) and transformer-based (SegFormer) segmentation architectures. The RescueNet dataset, a high-resolution post-disaster semantic segmentation benchmark, was adapted for detection by extracting bounding boxes from semantic masks via connected component analysis. Dataset adaptation was performed through separate pipelines for each annotation format (YOLO txt, torchvision JSON), and for segmentation, individual building crops were used, simulating the real-world cascaded pipeline scenario. At the segmentation stage, the task is formulated as binary building-footprint delineation within detected crops; per-building damage-level grading is performed by the subsequent classification stage of the cascade and is beyond the scope of this study. To ensure fair comparison, all models within each task were trained with identical hyperparameters, loss functions, and equivalent pretrained weights (COCO for detection, ImageNet for segmentation). Detection models were evaluated using standard COCO metrics (mAP@0.5, mAP@0.5:0.95) computed via pycocotools, while segmentation models were assessed using globally accumulated IoU, Dice, Precision, and Recall. Inference latency was additionally measured for all models. The results provide an evidence-based foundation for selecting architectures in a cascaded Detection-Segmentation-Classification pipeline for emergency response scenarios. The practical significance of the study lies in the unified evaluation framework that minimizes confounding factors such as different pretraining levels or optimizer configurations, enabling a substantially more direct comparison of architectural properties. This work is part of a broader project on developing an automated UAV-based building damage monitoring system.
Downloads
References
M. Rahnemoonfar, T. Chowdhury, and R. Murphy, “RescueNet: A High Resolution UAV Semantic Segmentation Dataset for Natural Disaster Damage Assessment,” Scientific Data, vol. 10, Art. no. 913, 2023, doi: 10.1038/s41597-023-02799-4.
S. Hafner, S. Gerard, J. Sullivan, and Y. Ban, “DisasterAdaptiveNet: A robust network for multi-hazard building damage detection from very-high-resolution satellite imagery,” International Journal of Applied Earth Observation and Geoinformation, vol. 143, Art. no. 104756, 2025, doi: 10.1016/j.jag.2025.104756.
S. Balovsyak and S. Stets, “Preprocessing of Object Images Before Their Detection Using YOLO Neural Network,” Security of Infocommunication Systems and Internet of Things, vol. 3, no. 2, Art. no. 02002, Dec. 2025, doi: 10.31861/sisiot2025.2.02002.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” in Proc. Advances in Neural Information Processing Systems (NIPS), 2015, pp. 91–99.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980–2988.
G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics YOLO,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” in Proc. Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015, pp. 234–241.
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation,” in Proc. European Conference on Computer Vision (ECCV), 2018, pp. 801–818.
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2021, pp. 12077–12090.
A. M. Braik and M. Koliou, “Automated building damage assessment and large-scale mapping by integrating satellite imagery, GIS, and deep learning,” Computer-Aided Civil and Infrastructure Engineering, vol. 39, no. 15, pp. 2389–2404, 2024, doi: 10.1111/mice.13197.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature Pyramid Networks for Object Detection,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2117–2125.
T. Adli, D. M. Bujaković, B. P. Bondžulić, M. Z. Laidouni, and M. S. Andrić, “Robustness of YOLO models for object detection in remote sensing images,” Journal of Electrical Engineering, vol. 76, no. 5, pp. 429–442, 2025, doi: 10.2478/jee-2025-0045.
G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, X. Xie, and J. Han, “Towards Large-Scale Small Object Detection: Survey and Benchmarks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 11, pp. 13467–13488, 2023, doi: 10.1109/TPAMI.2023.3290594.
H. Xia, J. Wu, J. Yao, H. Zhu, A. Gong, J. Yang, L. Hu, and F. Mo, “A Deep Learning Application for Building Damage Assessment Using Ultra-High-Resolution Remote Sensing Imagery in Turkey Earthquake,” International Journal of Disaster Risk Science, vol. 14, no. 6, pp. 947–962, 2023, doi: 10.1007/s13753-023-00526-6.
P. Iakubovskii, “Segmentation Models Pytorch,” 2019. [Online]. Available: https://github.com/qubvel-org/segmentation_models.pytorch
Published
Issue
Section
License
Copyright (c) 2026 Security of Infocommunication Systems and Internet of Things

This work is licensed under a Creative Commons Attribution 4.0 International License.









