Abstract
Level of detail 2 3-D modeling of buildings has rapidly developed, especially in digital twin and urban planning applications. However, the diversity of roof structures makes significant challenges for 3-D building modeling. While airborne images are a key data source for capturing optical roof features, they often struggle in low-contrast areas or shadowed conditions. LiDAR data can provide height information to compensate for the occluded areas. However, in the existing literature, particularly in existing deep learning models, the integration of these bimodal data has not been effectively studied for segmentation and modeling. This study proposes a novel multiobject instance segmentation framework that jointly detects all geometrical roof components required for 3-D building modeling, such as planes, inlines, and outlines, in a single bimodal network. Our method uses a shared ResNet-50 backbone with feature pyramid networks and introduces two custom attention mechanisms: intermodality attention blocks for fusing elevation and spectral features, and modality-specific attention blocks for refining modality-specific information. We also evaluate three fusion strategies, i.e., early, middle, and late, to optimize bimodal integration. Extracted lines and planes are vectorized, combined with height data from point clouds, and used to create 3-D planes and lines. The final building models and wireframes are constructed by intersecting these 3-D components. To the best of our knowledge, this is the first study to combine bimodal multiobject instance segmentation with attention mechanisms to jointly extract all roof components and create 3-D building models. Our study demonstrates that combining the optical and elevation features of each building improves the accuracy of building geometric component segmentation and 3-D building modeling. Experiments conducted with 1488 buildings show an F2-score of 0.8620, an F1-score of 0.8374, and an intersection over union of 0.844 for the segmentation step, and a root mean square error of around 1 m for the 3-D modeling.
Author supplied keywords
Cite
CITATION STYLE
Vostikolaei, F. S., & Jabari, S. (2025). Automated LoD2 Building Reconstruction Using Bimodal Segmentation. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 18, 23289–23305. https://doi.org/10.1109/JSTARS.2025.3602414
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.