Abstract
Automated building extraction using deep learning's semantic segmentation techniques on remote sensing imagery is increasingly important for mapping, urban planning, and environmental monitoring. However, this task presents considerable challenges, particularly when using Unmanned Aerial Vehicle (UAV) imagery from developing countries like Vietnam. The construction landscape in Vietnam is often characterized by non-homogeneous and non-standard building designs, as well as the coexistence of slums and modern urban areas. These real-world issues, combined with a lack of annotated data, frequently lead to model overfitting and reduced generalization. This study aims to fill a critical knowledge gap by systematically evaluating six widely used deep learning architectures: DeepLabV3+, Unet, Unet++, Feature Pyramid Network (FPN), Pyramid Attention Network (PAN), and SegFormer, specifically for building extraction tasks. We utilize Intersection over Union (IoU) scores and analyze overfitting behavior to compare the performance of these models on a benchmark WHU Building Dataset and a challenging proprietary Vietnamese UAV Building Dataset. Our thorough analysis quantifies the individual and combined effects of transfer learning and data augmentation on model performance. Our findings indicate that both transfer learning and data augmentation are essential, consistently resulting in higher IoU scores improving by 10-15% with transfer learning—and significantly reducing overfitting (from as much as 30% down to 3-7% on the Vietnamese dataset). Furthermore, architectural efficiency is more critical than the raw parameter count. Models like DeepLabV3+ and SegFormer achieved top-tier performance, with IoU scores of 0.858 and 0.855, respectively, on the Vietnamese dataset while maintaining a moderate parameter count of 11-12 million. The Vietnamese dataset proved to be inherently more challenging than the WHU benchmark, highlighting the need for robust training strategies for effective generalization in complex, real-world environments. This research provides recommendations for selecting appropriate model architectures and training strategies to ensure accurate and reliable building extraction under difficult conditions.
Author supplied keywords
Cite
CITATION STYLE
Trung, D. P., & Ngoc, H. P. (2025). Deep Learning Architectures for Building Extraction from UAV Imagery: A Comparative Analysis of Transfer Learning and Data Augmentation Strategies. Inzynieria Mineralna, 1(2), 609–624. https://doi.org/10.29227/IM-2025-02-50
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.