Remote sensing image semantic segmentation method combining cosine annealing with atrous convolution

  • role: First author第一作者
  • Affiliation:

    School of Civil and Hydraulic Engineering, Huazhong University of Science and Technology, Wuhan 430074, China

  • Email:u201715748@hust.edu.cn
  • Introduction:E-mail u201715748@hust.edu.cn
TANG Zhenchao1,  
  • Affiliation:

    Yellow River Survey, Planning, Design and Research Institute Co., Ltd, Zhengzhou 450003, China

WEI Wei2,  
  • Affiliation:

    School of Water Conservancy and Environment, Zhengzhou University, Zhengzhou 450001, China

LUO Weiran3,  
  • Affiliation:

    Yellow River Survey, Planning, Design and Research Institute Co., Ltd, Zhengzhou 450003, China

HU Jie2,  
  • role: Corresponding author通信作者
  • Affiliation:

    School of Civil and Hydraulic Engineering, Huazhong University of Science and Technology, Wuhan 430074, China

  • Email:zhangdongying@hust.edu.cn
  • Introduction:E-mailzhangdongying@hust.edu.cn
ZHANG Dongying1*

résumé

This study aims to capture the rich context information and multiscale feature information in remote sensing images, improve the integrated model strategy, and enhance the accuracy of semantic segmentation. Thus, this study proposes a high-resolution remote sensing image semantic segmentation method using cosine annealing with increasing period and multiscale atrous convolution.The multiscale parallel atrous convolution helps the network capture context information in a larger range and improves the ability of the network to recognize multiscale objects without increasing parameters. The method in this study uses the atrous convolution while discarding the pooling operation to maintain the spatial resolution. Meanwhile, the method adopts the fully connected conditional random field to add spatial and edge context information for making up for part of the position information missed by the atrous convolution. As a result, the outline of extraction objects by semantic segmentation fits the ground truth better. Moreover, the cosine annealing strategy with increasing period is introduced to adjust the learning rate and obtain a suitable number of local optimal solutions. We integrate the local optimal solutions in the method to further improve the pixel classification ability of the network.The overall accuracy and kappa coefficient of the proposed model, which are 86.6% and 81.8%, respectively, are better than those of the current advanced semantic segmentation models.The experimental results performed on the Gaofen image dataset show that the fusion of image context information and multiscale feature information can effectively identify objects with complex structures. Moreover, the model coupled with the period-increasing cosine annealing strategy could obtain better semantic segmentation accuracy than and less inference time than that coupled with the equal-period cosine annealing strategy.

mots-clés

high-resolution remote sensing image;semantic segmentation;cosine annealing with increasing period;multi-scale parallel atrous convolution;target extraction;in-context learning;conditional random field;multi-scale learning

References

  1. 1.
    Anthimopoulos M, Christodoulidis S, Ebner L, Geiser T, Christe A and Mougiakakou S. 2019. Semantic segmentation of pathological lung tissue with dilated fully convolutional networks. IEEE Journal of Biomedical and Health Informatics, 23(2): 714-722
  2. 2.
    Badrinarayanan V, Kendall A and Cipolla R. 2017. SegNet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12): 2481-2495
  3. 3.
    Chen L C, Papandreou G, Kokkinos I, Murphy K and Yuille A L. 2015. Semantic image segmentation with deep convolutional nets and fully connected CRFs//3rd International Conference on Learning Representations. San Diego: ICLR
  4. 4.
    Chen L C, Papandreou G, Kokkinos I, Murphy K and Yuille A L. 2018. DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4): 834-848
  5. 5.
    Chen L C, Papandreou G, Schroff F and Adam H. 2017. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706. 05587
  6. 6.
    Deng J, Dong W, Socher R, Li L J, Li K and Fei-Fei L. 2009. ImageNet: a large-scale hierarchical image database//2009 IEEE Conference on Computer Vision and Pattern Recognition. Miami: IEEE: 248-255
  7. 7.
    Dumoulin V and Visin F. 2016. A guide to convolution arithmetic for deep learning. arXiv preprint arXiv:1603. 07285
  8. 8.
    Garcia-Garcia A, Orts-Escolano S, Oprea S, Villena-Martinez V and Garcia-Rodriguez J. 2017. A review on deep learning techniques applied to semantic segmentation. arXiv preprint arXiv:1704. 06857
  9. 9.
    He K M, Zhang X Y, Ren S Q and Sun J. 2016. Deep residual learning for image recognition//Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 770-778
  10. 10.
    Hinton G, Vinyals O and Dean J. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503. 02531
  11. 11.
    Huang G, Li Y X, Pleiss G, Liu Z, Hopcroft J E and Weinberger K Q. 2017. Snapshot ensembles: train 1, get m for free//5th International Conference on Learning Representations. Toulon: ICLR
  12. 12.
    Ioffe S and Szegedy C. 2015. Batch normalization: accelerating deep network training by reducing internal covariate shift//Proceedings of the 32nd International Conference on International Conference on Machine Learning. Lille: JMLR.org: 448-456
  13. 13.
    Kamann C and Rother C. 2020. Benchmarking the robustness of semantic segmentation models//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE: 8825-8835
  14. 14.
    Kingma D P and Ba J. 2015. Adam: a method for stochastic optimization//3rd International Conference on Learning Representations. San Diego: ICLR
  15. 15.
    Krähenbühl P and Koltun V. 2011. Efficient inference in fully connected CRFs with Gaussian edge potentials//Proceedings of the 24th International Conference on Neural Information Processing Systems. Granada: Curran Associates Inc.: 109-117
  16. 16.
    Li Y, Xiao C J, Zhang H Q, Li X J and Chen J. 2020. Remote sensing image semantic segmentation using deep fusion convolutional networks and conditional random field. Remote Sensing for Natural Resources, 32(3): 15-22
  17. 17.
    Long J, Shelhamer E and Darrell T. 2015. Fully convolutional networks for semantic segmentation//Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE: 3431-3440
  18. 18.
    Loshchilov I and Hutter F. 2017. SGDR: stochastic gradient descent with warm restarts//5th International Conference on Learning Representations. Toulon: ICLR
  19. 19.
    Polino A, Pascanu R and Alistarh D. 2018. Model compression via distillation and quantization//6th International Conference on Learning Representations. Vancouver: ICLR
  20. 20.
    Ronneberger O, Fischer P and Brox T. 2015. U-Net: convolutional networks for biomedical image segmentation//Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention. Munich: Springer: 234-241
  21. 21.
    Simonyan K and Zisserman A. 2015. Very deep convolutional networks for large-scale image recognition//3rd International Conference on Learning Representations. San Diego: ICLR
  22. 22.
    Sun W W and Wang R S. 2018. Fully convolutional networks for semantic segmentation of very high resolution remotely sensed images combined with DSM. IEEE Geoscience and Remote Sensing Letters, 15(3): 474-478
  23. 23.
    Teichmann M and Cipolla R. 2019. Convolutional CRFs for semantic segmentation//30th British Machine Vision Conference 2019. Cardiff: BMVC: 142
  24. 24.
    Tong X Y, Xia G S, Lu Q K, Shen H F, Li S Y, You S C and Zhang L P. 2020. Land-cover classification with high-resolution remote sensing images using transferable deep models. Remote Sensing of Environment, 237: 111322
  25. 25.
    Wang P Q, Chen P F, Yuan Y, Liu D, Huang Z H, Hou X D and Cottrell G. 2018. Understanding convolution for semantic segmentation//2018 IEEE Winter Conference on Applications of Computer Vision (WACV). Lake Tahoe: IEEE: 1451-1460
  26. 26.
    Wang Z W, Wang Z P, You S C, Lei F, Cao L and Yang K J. 2020. Landsat image glacier extraction based on context semantic segmentation network. Acta Geodaetica et Cartographica Sinica, 49(12): 1575-1582
  27. 27.
    Yu F and Koltun V. 2016. Multi-scale context aggregation by dilated convolutions//4th International Conference on Learning Representations. San Juan: ICLR
  28. 28.
    Zeiler M D. 2012. ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701
  29. 29.
    Zhao H S, Shi J P, Qi X J, Wang X G and Jia J Y. 2017. Pyramid scene parsing network//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE: 6230-6239
  30. 30.
    Zhao Q H, Xie K L, Wang G H and Li Y. 2020. Land cover classification of polarimetric SAR with fully convolution network and conditional random field. Acta Geodaetica et Cartographica Sinica, 49(1): 65-78
  31. 31.
    Zhou P C, Cheng G, Yao X W and Han J W. 2021. Machine learning paradigms in high-resolution remote sensing image interpretation. National Remote Sensing Bulletin, 25(1): 182-197
  32. 32.
    Zuo Z C, Zhang W and Zhang D Y. 2020. A remote sensing image semantic segmentation method by combining deformable convolution with conditional random fields. Journal of Geodesy and Geoinformation Science, 3(3): 39-49

Lire l'article complet

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website