Res_ASPP_UNet++: Building an extraction network from remote sensing imagery combining depthwise separable convolution with atrous spatial pyramid pooling

  • role: First author第一作者
  • Affiliation:

    School of Land Resources Engineering, Kunming University of Science and Technology, Kunming 650093, China

  • Email:2064327704@qq.com
  • Introduction:E-mail 2064327704@qq.com
LYU Shaoyun,  
  • role: Corresponding author通信作者
  • Affiliation:

    School of Land Resources Engineering, Kunming University of Science and Technology, Kunming 650093, China

  • Email:ljtwcx@163.com
  • Introduction:E-mail ljtwcx@163.com
LI Jiatian*,  
  • Affiliation:

    School of Land Resources Engineering, Kunming University of Science and Technology, Kunming 650093, China

A Xiaohui,  
  • Affiliation:

    School of Land Resources Engineering, Kunming University of Science and Technology, Kunming 650093, China

YANG Chao,  
  • Affiliation:

    School of Land Resources Engineering, Kunming University of Science and Technology, Kunming 650093, China

YANG Ruchun,  
  • Affiliation:

    School of Land Resources Engineering, Kunming University of Science and Technology, Kunming 650093, China

SHANG Xiaomei

Resümee

Building extraction from remote sensing imagery is an important research direction for the interpretation of remote sensing imagery. UNet++ is constructed from U-Net by connecting the decoders, resulting in densely connected skip connections, enabling dense feature propagation along skip connections and thus more flexible feature fusion at the decoder nodes. However, the traditional standard convolution in the encoder fails to fully capture the multiscale features of remote sensing imagery because of the single path for semantic feature extraction, which affects the segmentation performance of the network to some extent. To address this problem, we propose a building extraction method to improve the accuracy of building extraction from remote sensing imagery.On the basis of UNet++, whose backbone is a deep residual network, we propose a building extraction network by replacing the standard convolution and max pooling in the encoder with depthwise separable convolution and applying an atrous spatial pyramid pooling structure (ASPP) to the end of the encoder. The network is referred to as the residual atrous spatial pyramid network (Res_ASPP_UNet++). On the basis of using dense and short connections to reduce the semantic gap between the encoder and the decoder, the Res_ASPP_UNet++ architecture uses multiscale ASPP made of several atrous convolutions with different sampling rates to sample the image in parallel to enrich semantic information by expanding the field of view and applies image-level features to encode the global context, avoiding segmentation errors caused by local features and improving the target segmentation performance.Experiments are conducted to validate the effectiveness of the proposed methodology. We compare the frequently used semantic segmentation networks with the Res_ASPP_UNet++ architecture using the WHU and Massachusetts datasets as data sources and take intersection over union (IoU), accuracy, precision and F1-score as evaluation indexes to evaluate the accuracy of building extraction. The experimental results are as follows: (1) The integrity of buildings extracted by Res_ASPP_UNet++ network is better than other segmentation networks on the whole, the boundary of the segmentation result is smoother and more accurate, and the result has less segmentation noise; (2) Compared with the compared semantic segmentation networks, the number of parameters of Res_ASPP_UNET ++ network is greatly compressed after improvement, the segmentation accuracy of the method is better on the whole, and it has great improvement to the UNET++ network through quantitative analysis; (3) Res_ASPP_UNET++ network has higher IoU index values than the existing building extraction algorithms listed in this paper; (4) Comparied with the proposed method without ASPP module, the value of each evaluation indexes of Res_ASPP_UNet++ with ASPP module is higher and the whole accuracy of the network with depthwise separable convolution is slightly higher than that of the network with max pooling; (5) Res_ASPP_UNet++ also extract the buildings on the Massachusetts dataset effectively and is better than other segmentation networks on the whole.The following conclusions can be drawn from the experimental results. Multiscale ASPP and depthwise separable convolution can improve the ability to express the detailed features of the model and thus effectively improve the building extraction performance from remote sensing imagery on the basis of greatly compressing the model parameters. In addition, the Res_ASPP_UNet++ architecture is robust to buildings of different scales and types and has strong generalization ability on datasets of different resolutions and sources. In future work, we will improve the segmentation for areas with a complex distribution of buildings and overcome the missing segmentation of small buildings.

Schlüsselwort

remote sensing imagery;building extraction;UNet++;depthwise separable convolution;deep residual structure;atrous spatial pyramid pooling

References

  1. 1.
    Badrinarayanan V, Kendall A and Cipolla R. 2017. SegNet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12): 2481-2495
  2. 2.
    Berman M, Triki A R and Blaschko M B. 2018. The lovász-softmax loss: a tractable surrogate for the optimization of the intersection-over-union measure in neural networks//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE: 4413-4421
  3. 3.
    Chen K Q, Gao X, Yan M L, Zhang Y and Sun X. 2020. Building extraction in pixel level from aerial imagery with a deep encoder-decoder network. Journal of Remote Sensing, 24(9): 1134-1142
  4. 4.
    Chen L C, Papandreou G, Schroff F and Adam H. 2017. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv: 1706.05587
  5. 5.
    Chen L C, Zhu Y K, Papandreou G, Schroff F and Adam H. 2018. Encoder-decoder with atrous separable convolution for semantic image segmentation//Proceedings of the 15th European Conference on Computer Vision. Munich: Springer: 801-818
  6. 6.
    Chollet F. 2017. Xception: deep learning with depthwise separable convolutions//Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: IEEE: 1800-1807
  7. 7.
    He K M, Zhang X Y, Ren S Q and Sun J. 2016. Deep residual learning for image recognition//Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas, NV, USA: IEEE: 770-778
  8. 8.
    Howard A G, Zhu M L, Chen B, Kalenichenko D, Wang W J, Weyand T, Andreetto M and Adam H. 2017. MobileNets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv: 1704.04861
  9. 9.
    Huang X and Zhang L P. 2011. A multidirectional and multiscale morphological index for automatic building extraction from multispectral GeoEye-1 imagery. Photogrammetric Engineering and Remote Sensing, 77(7): 721-732
  10. 10.
    Ji S P, Tian S Q and Zhang C. 2020. Urban land cover classification and change detection using fully atrous convolutional neural network. Geomatics and Information Science of Wuhan University, 45(2): 233-241
  11. 11.
    Kang W C, Xiang Y M, Wang F and You H J. 2019. EU-Net: an efficient fully convolutional network for building extraction from optical remote sensing images. Remote Sensing, 11(23): 2813
  12. 12.
    Krizhevsky A, Sutskever I and Hinton G E. 2012. ImageNet classification with deep convolutional neural networks//Proceedings of the 25th International Conference on Neural Information Processing Systems. Lake Tahoe, Nevada: Curran Associates Inc.: 1097-1105
  13. 13.
    Lin G S, Milan A, Shen C H and Reid I. 2017. RefineNet: multi-path refinement networks for high-resolution semantic segmentation//Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: IEEE: 5168-5177
  14. 14.
    Lin J B, Jing W P, Song H B and Chen G S. 2019. ESFNet: efficient network for building extraction from high-resolution aerial images. IEEE Access, 7: 54285-54294 [DOT: ]
  15. 15.
    Mnih V. 2013. Machine Learning for Aerial Image Labeling. Toronto: University of Toronto
  16. 16.
    Pesaresi M, Gerhardinger A and Kayitakire F. 2008. A robust built-up area presence index by anisotropic rotation-invariant textural measure. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 1(3): 180-192
  17. 17.
    Ronneberger O, Fischer P and Brox T. 2015. U-Net: convolutional networks for biomedical image segmentation//Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention. Munich, Germany: Springer: 234-241
  18. 18.
    Saito S and Aoki Y. 2015. Building and road detection from large aerial imagery//Proceedings of SPIE 9405, Image Processing: Machine Vision Applications VIII. San Francisco: SPIE
  19. 19.
    Shelhamer E, Long J and Darrell T. 2017. Fully convolutional networks for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4): 640-651
  20. 20.
    Sherrah J. 2016. Fully convolutional networks for dense semantic labelling of high-resolution aerial imagery. arXiv preprint arXiv: 1606.02585
  21. 21.
    Simonyan K and Zisserman A. 2015. Very deep convolutional networks for large-scale image recognition//Proceedings of 3rd International Conference on Learning Representations. San Diego: ICLR: 463-476
  22. 22.
    Szegedy C, Liu W, Jia Y Q, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V and Rabinovich A. 2015. Going deeper with convolutions//Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston, MA: IEEE: 1-9 [DOT: 10.1109/cvpr.2015.7298594]
  23. 23.
    Tan Q L. 2010. Urban building extraction from VHR multi-spectral images using object-based classification. Acta Geodaetica et Cartographica Sinica, 39(6): 618-623
  24. 24.
    Tao W B, Liu J and Tian J W. 2003. A new approach to extract rectangle building automatically from aerial images. Chinese Journal of Computers, 26(7): 866-873
  25. 25.
    Wang J, Qin Q M, Ye X, Wang J H, Qin X B and Yang X C. 2016. A survey of building extraction methods from optical high resolution remote sensing imagery. Remote Sensing Technology and Application, 31(4): 653-662
  26. 26.
    Wang W, Yang X C, Qin X B, Ye X and Qin Q M. 2015. An efficient approach for automatic rectangular building extraction from very high resolution optical satellite imagery. IEEE Geoscience and Remote Sensing Letters, 12(3): 487-491
  27. 27.
    Wu G M, Chen Q, Shibasaki R, Guo Z L, Shao X W and Xu Y W. 2018. High precision building detection from aerial imagery using a U-net like convolutional architecture. Acta Geodaetica et Cartographica Sinica, 47(6): 864-872
  28. 28.
    Wu G M, Shao X W, Guo Z L, Chen Q, Yuan W, Shi X D, Xu Y W and Shibasaki R. 2018. Automatic building segmentation of aerial imagery using multi-constraint fully convolutional networks. Remote Sensing, 10(3): 407
  29. 29.
    Wu W, Luo J C, Shen Z F and Zhu Z W. 2012. Building extraction from high resolution remote sensing imagery based on spatial-spectral method. Geomatics and Information Science of Wuhan University, 37(7): 800-805
  30. 30.
    You Y F, Wang S Y, Wang B, Ma Y X, Shen M, Liu W H and Xiao L. 2019. Study on hierarchical building extraction from high resolution remote sensing imagery. Journal of Remote Sensing, 23(1): 125-136
  31. 31.
    Zhang C S, Ge Y W and Jiang X. 2020. High-resolution remote sensing image building extraction based on sparsely constrained SegNet. Journal of Xi’an University of Science and Technology, 40(3): 441-448
  32. 32.
    Zhang Z X and Wang Y H. 2019. JointNet: a common neural network for road and building extraction. Remote Sensing, 11(6): 696
  33. 33.
    Zhao W Z, Guo Z, Yue J, Zhang X Y and Luo L Q. 2015. On combining multiscale deep learning features for the classification of hyperspectral remote sensing imagery. International Journal of Remote Sensing, 36(13): 3368-3379
  34. 34.
    Zhou D J, Wang G Z, He G J, Long T F, Yin R Y, Zhang Z M, Chen S B and Luo B. 2020. Robust building extraction for high spatial resolution remote sensing images with self-attention network. Sensors, 20(24): 7241
  35. 35.
    Zhou Z W, Siddiquee M, Tajbakhsh N and Liang J M. 2018. UNet++: a nested U-Net architecture for medical image segmentation//Proceedings of the 8th International Workshop on Deep Learning in Medical Image Analysis. Granada: Springer: 3-11

Lesen Sie die ganze Passage

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website