A dual-attention capsule network for building extraction from high-resolution remote sensing imagery

  • No information about the author is available
Zhengsen XU,  
  • No information about the author is available
Haiyan GUAN,  
  • No information about the author is available
Daifeng PENG,  
  • No information about the author is available
Yongtao YU,  
  • No information about the author is available
Xiangda LEI,  
  • No information about the author is available
Haohao ZHAO

ملخص

Automatic extraction of buildings from high-resolution remote sensing images is greatly important in disaster prevention and mitigation, disaster loss estimation, urban planning, and topographic map making. With the advancement of optical remote sensing techniques in image resolutions and qualities, remote sensing images have provided an important data source for assisting the rapid updating of building footprint database. Despite the large number of algorithms proposed with enhanced performance, fulfilling highly accurate and fully automated extraction of buildings from remote sensing images is still difficult due to the considerable challenging scenarios of buildings, such as color diversities, topology variations, occlusions, and shadow covers. Thus, exploiting advanced and high-performance techniques to further improve the accuracy and automation level of building extraction is greatly meaningful and urgently required by a large variety of applications.To overcome the issues of strong variability and weak homogeneity of traditional convolutional neural networks, we propose a novel dual-attention capsule encoder–decoder network DA-CapsNet for extracting buildings. In this network, a deep capsule encoder–decoder network, along with the channel-spatial attention blocks, is developed to enhance the capability of extracting high-level feature information from very high resolution remote sensed images. Thus, this model has the ability to extract buildings covered by shadows and discriminate buildings from non-building impervious surfaces. Specifically, we initially employ a deep capsule encoder–decoder network to extract and fuse multiscale building capsule features, resulting in a high-quality building feature representation. Moreover, spatial attention and channel attention modules are designed to further rectify and enhance the captured contextual information to obtain a competitive performance in processing buildings in the diverse challenging scenarios. The contributions include the following: (1) the deep capsule encoder–decoder network is designed to generate a high-quality feature representation; (2) the channel and spatial feature attention modules are designed to highlight channel-wise salient features and focus on class-specific spatial features.The proposed DA-CapsNet was evaluated on three datasets: one Google Building Dataset and two publicly-available datasets (Wuhan and Massachusetts). The experimental results achieved a competitive performance with an average precision, recall, and F1-score of 92.15%, 92.07%, and 92.18%, respectively, in handling buildings of varying challenging scenarios. Considering the overall accuracy of F1-score, the DA-CapsNet achieved the values of 92.70%, 94.01%, and 89.84% for Google, WUH, and MA datasets, respectively. Comparative studies also confirmed the robust applicability and superior performance of the DA-CapsNet in building extraction tasks.

مفهوم

building extraction;deep learning;channel feature attention;spatial feature attention;encoder-decoder network;capsule network

قراءة النص الكامل

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website