Feature attention pyramid-based remote sensing image object detection method

  • role: First author第一作者
  • Affiliation:

    School of Computer Science, Shaanxi Normal University, Xi'an 710119, China

  • Email:wangxili@snnu.edu.cn
  • Introduction:西E-mail wangxili@snnu.edu.cn
WANG Xili,  
  • Affiliation:

    School of Computer Science, Shaanxi Normal University, Xi'an 710119, China

LIANG Zhengyin,  
  • Affiliation:

    School of Computer Science, Shaanxi Normal University, Xi'an 710119, China

LIU Tao

Resümee

The characteristics of remote sensing images, such as complex scenes, different sizes of targets and unbalanced distributions, increase the difficulty of target detection. However, feature pyramids that are suitable for detecting targets of different scales do not consider the importance of different feature maps when fusing the feature maps, let alone emphasize the features of target areas. For this purpose, this paper proposes a feature attention pyramid-based remote sensing image object detection method (namely, the feature attention pyramid network, FAPNet).First, the feature maps of different depths are fused by channel concatenation, and the features of different sized receptive fields are provided for the feature maps used for detection. The channel attention mechanism is used to recalibrate the fused feature maps in the channel dimension. The feature maps from different depths are adaptively adjusted according to the scale of the object to be detected to strengthen the feature that matches highly between the size of the receptive field and the object to be detected and weaken the feature with a low degree of matching. Second, the weakly supervised attention module uses the superimposed atrous spatial pyramid pooling structure and convolutional segmentation module to model spatial attention weights to adjust the feature distribution of the feature map used for prediction, strengthen the object area feature, and weaken the background area feature, which further improves the performance of object detection methods.The experimental results show that compared with RetinaNet, the proposed method improves the accuracy (AP) for car targets by 3.41% and 2.26% on the UCAS-AOD dataset and RSOD dataset, respectively, achieves better AP results on each target for multiclass targets and is superior to other comparative object detection methods on the mAP indicator for multitargets.A feature attention pyramid-based remote sensing image object detection method is proposed in this paper. Its contribution lies in the designed feature attention pyramid module and weakly supervised attention module. With the new modules, the proposed method can extract target features more accurately in complex scenes with targets of different sizes by channel attention and spatial attention, thus improving the performance of detection. The experimental results show that the proposed method is superior to the RetinaNet and FAN methods and is more suitable for remote sensing image object detection tasks with complex scenes and multiscale targets.

Schlüsselwort

remote sensing image;object detection;weakly supervised segmentation;attention mechanism;feature pyramid

References

  1. 1.
    Chen L C, Papandreou G, Kokkinos I, Murphy K and Yuille A L. 2018. DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4): 834-848
  2. 2.
    Dai J F, Li Y, He K M and Sun J. 2016. R-FCN: object detection via region-based fully convolutional networks//Proceedings of the 30th International Conference on Neural Information Processing Systems. Barcelona, Spain: Curran Associates Inc.: 379-387
  3. 3.
    Dalal N and Triggs B. 2005. Histograms of oriented gradients for human detection//Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. San Diego, CA, USA: IEEE: 886-893
  4. 4.
    Fu C Y, Liu W, Ranga A, Tyagi A and Berg A C. 2017. DSSD: deconvolutional single shot detector. arXiv preprint arXiv: 1701.06659
  5. 5.
    He K M, Zhang X Y, Ren S Q and Sun J. 2016. Deep residual learning for image recognition//Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas, NV, USA: IEEE: 770-778
  6. 6.
    Hu J, Shen L, Albanie S, Sun G and Wu E H. 2020. Squeeze-and-excitation networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(8): 2011-2023
  7. 7.
    Hu P Y and Ramanan D. 2017. Finding tiny faces//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: 1522-1530
  8. 8.
    Huang G, Liu Z, Van Der Maaten L and Weinberger K Q. 2017. Densely connected convolutional networks//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: IEEE: 2261-2269
  9. 9.
    Lin T Y, Dollár P, Girshick R, He K M, Hariharan B and Belongie S. 2017. Feature pyramid networks for object detection//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: IEEE: 936-944
  10. 10.
    Lin T Y, Goyal P, Girshick R, He K M and Dollár P. 2020. Focal loss for dense object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2): 318-327
  11. 11.
    Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C Y and Berg A C. 2016. SSD: single shot multiBox detector//Proceedings of the 2016 14th European Conference on Computer Vision. Amsterdam: Springer: 21-37
  12. 12.
    Long Y, Gong Y P, Xiao Z E and Liu Q. 2017. Accurate object localization in remote sensing images based on convolutional neural networks. IEEE Transactions on Geoscience and Remote Sensing, 55(5): 2486-2498
  13. 13.
    Redmon J and Farhad A. 2017. YOLO9000: better, faster, stronger//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: IEEE: 6517-6525
  14. 14.
    Ren S Q, He K M, Girshick R and Sun J. 2017. Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6): 1137-1149
  15. 15.
    Singh B and Davis L S. 2018. An analysis of scale invariance in object detection-SNIP//Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, UT, USA: IEEE: 3578-3587
  16. 16.
    Tayara H and Chong K T. 2018. Object detection in very high-resolution aerial images using one-stage densely connected feature pyramid network. Sensors, 18(10): 3341
  17. 17.
    Wang F, Jiang M Q, Qian C, Yang S, Li C, Zhang H G, Wang X G and Tang X O. 2017a. Residual attention network for image classification//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, HI, USA: IEEE: 6450-6458
  18. 18.
    Wang J F, Yuan Y and Yu G. 2017b. Face attention network: an effective face detector for the occluded faces. arxiv preprint arXiv: 1711.07246
  19. 19.
    Woo S, Park J, Lee J Y and Kweon I S. 2018. CBAM: convolutional block attention module//Proceedings of 2018 the 15th European Conference on Computer Vision. Munich: Springer: 3-19
  20. 20.
    Yang A P, Lu L Y and Ji Z. 2020. Multi-feature concatenation network for object detection. Journal of Tianjin University (Natural Science and Engineering), 53(6): 647-652
  21. 21.
    Yu Y, Hua A, He X J, Yu S H, Zhong X and Zhu R F. 2020. Attention-based feature pyramid networks for ship detection of optical remote sensing image. Journal of Remote Sensing, 24(2): 107-115
  22. 22.
    Zhou K B, Zhang Z X, Gao C X and Liu J. 2021. Rotated feature network for multiorientation object detection of remote-sensing images. IEEE Geoscience and Remote Sensing Letters, 18(1): 33-37
  23. 23.
    Zhou P, Ni B B, Geng C, Hu J G and Xu Y. 2018. Scale-transferrable object detection//Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, UT, USA: IEEE: 528-537
  24. 24.
    Zhu H G, Chen X G, Dai W Q, Fu K, Ye Q X and Jiao J B. 2015. Orientation robust object detection in aerial images using deep convolutional neural network//Proceedings of the 2015 IEEE International Conference on Image Processing. Quebec City, QC, Canada: IEEE: 3735-3739

Lesen Sie die ganze Passage

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website