Application of an improved CenterNet in remote sensing images object detection

  • role: First author第一作者
  • Affiliation:

    State Key Laboratory of Complex Electromagnetic Environment Effects on Electronics and Information System (CEMEE), Luoyang 471000, China

  • Email:tzz14@nudt.edu.cn
  • Introduction:E-mailtzz14@nudt.edu.cn
TIAN Zhuangzhuang1,  
  • Affiliation:

    State Key Laboratory of Complex Electromagnetic Environment Effects on Electronics and Information System (CEMEE), Luoyang 471000, China

ZHANG Hengwei1,  
  • Affiliation:

    State Key Laboratory of Complex Electromagnetic Environment Effects on Electronics and Information System (CEMEE), Luoyang 471000, China

WANG Kun1,  
  • Affiliation:

    College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China

LIU Shengqi2,  
  • Affiliation:

    State Key Laboratory of Complex Electromagnetic Environment Effects on Electronics and Information System (CEMEE), Luoyang 471000, China

ZOU Qianjin1,  
  • Affiliation:

    State Key Laboratory of Complex Electromagnetic Environment Effects on Electronics and Information System (CEMEE), Luoyang 471000, China

ZHAO Zhen1,  
  • Affiliation:

    State Key Laboratory of Complex Electromagnetic Environment Effects on Electronics and Information System (CEMEE), Luoyang 471000, China

CHEN Yubin1

resumen

Nowadays, object detection methods based on deep learning are widely used in the interpretation of remote sensing images. The anchor-based methods usually need to design the anchor boxes first, which requires more detection steps and time cost. This study proposed an object detection method of remote sensing images based on the improved CenterNet. The method can simplify the object detection process and improve efficiency.The CenterNet uses a fully convolutional network to directly predict the heat map of the center points, widths, and heights of the corresponding objects, and the position offsets of the center points. The heat maps are used to generate the rough positions of the objects, and the offsets can fine-tune the positions to make them more accurate. The widths and heights further constitute the shape of the object boxes. The different heat maps decide the object categories. On the basis of CenterNet, the proposed method first adopts the ResNet with transposed convolution as the backbone network. The transposed convolution can expand the output feature maps, and ResNet can reduce the number of parameters in the backbone network compared with the Hourglass network. Second, the proposed method defines the length of Gaussian kernel under three limit conditions between the predicted and real boxes in CenterNet. The Gaussian kernel is applied to generate the heat map label, which is used for network training. Finally, the multi-head attention mechanism is introduced into the backbone network to learn the importance of each element in the feature maps. The weights assigned to the elements reflect their effectiveness, which makes the effective features concentrate in the regions of the object key points as much as possible.The experiments use mean Average Precision (mAP) to evaluate the object detection results on the multiple categories. All the experiments are conducted at the DIOR dataset. The results show that the CenterNet using the ResNet with transposed convolution is 1.4% higher than that using the Hourglass. The proposed calculation of the length of the Gaussian kernel can increase the mAP by 1.1%. The addition of attention mechanism can further improve the mAP by 1.5%. At the same time, the proposed method reduces the time cost by 31.9% compared with the conventional method.The experimental results show that the proposed method can improve detection accuracy without sacrificing the detection speed. The ablation experiments of different parts also show that the ResNet with transposed convolution, the designed calculation method of the length of the Gaussian kernel, and the attention mechanism can effectively improve the mAP. The comparison with other methods also proves that the proposed method is practical.

palabra clave

remote sensing image;object detection;deep learning;CenterNet;attention mechanism

References

  1. 1.
    Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A and Zagoruyko S. 2020. End-to-end object detection with transformers. arXiv: 2005.12872
  2. 2.
    Chen Q, Wang Y M, Yang T, Zhang X Y, Cheng J and Sun J. 2021. You only look one-level feature//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE: 13034-13043
  3. 3.
    Chen R, Liu Y, Zhang M D, Liu S, Yu B and Tai Y W. 2020. Dive deeper into box for object detection//16th European Conference on Computer Vision. Glasgow: Springer: 412-428
  4. 4.
    Cheng G and Han J W. 2016. A survey on object detection in optical remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing, 117: 11-28
  5. 5.
    Dai Z G, Cai B L, Lin Y G and Chen J Y. 2021. UP-DETR: unsupervised pre-training for object detection with transformers//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE: 1601-1610
  6. 6.
    Girshick R, Donahue J, Darrell T and Malik J. 2016. Region-based convolutional networks for accurate object detection and segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(1): 142-158
  7. 7.
    He K M, Zhang X Y, Ren S Q and Sun J. 2016. Deep residual learning for image recognition//2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 770-778
  8. 8.
    Hong M B, Li S W, Yang Y C, Zhu F Y, Zhao Q J and Lu L. 2022. SSPNet: scale selection pyramid network for tiny person detection from UAV images. IEEE Geoscience and Remote Sensing Letters, 19: 8018505
  9. 9.
    Kong T, Sun F C, Liu H P, Jiang Y N, Li L and Shi J B. 2020. FoveaBox: beyound anchor-based object detection. IEEE Transactions on Image Processing, 29: 7389-7398
  10. 10.
    Law H and Deng J. 2018. CornerNet: detecting objects as paired keypoints//15th European Conference on Computer Vision. Munich: Springer: 765-781
  11. 11.
    Li K, Wan G, Cheng G, Meng L Q and Han J W. 2020. Object detection in optical remote sensing images: a survey and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing, 159: 296-307
  12. 12.
    Lin T Y, Dollár P, Girshick R, He K M, Hariharan B and Belongie S. 2017a. Feature pyramid networks for object detection//2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE: 936-944
  13. 13.
    Lin T Y, Goyal P, Girshick R, He K M and Dolár P. 2017b. Focal loss for dense object detection//2017 IEEE International Conference on Computer Vision. Venice: IEEE: 2999-3007
  14. 14.
    Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C Y and Berg A C. 2016. SSD: single shot MultiBox detector//14 European Conference on Computer Vision. Amsterdam: Springer: 21-37
  15. 15.
    Newell A, Yang K Y and Deng J. 2016. Stacked hourglass networks for human pose estimation//14th European Conference on Computer Vision. Amsterdam: Springer: 483-499
  16. 16.
    Ren S Q, He K M, Girshick R and Sun J. 2015. Faster R-CNN: towards real-time object detection with region proposal networks//Proceedings of the 28th International Conference on Neural Information Processing Systems. Montreal: MIT Press: 91-99
  17. 17.
    Tian Z, Shen C H, Chen H and He T. 2019. FCOS: fully convolutional one-stage object detection//2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE: 9626-9635
  18. 18.
    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser Ł and Polosukhin I. 2017. Attention is all you need//Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach: Curran Associates Inc.: 6000-6010
  19. 19.
    Wang C, Bai X, Wang S, Zhou J and Ren P. 2019a. Multiscale visual attention networks for object detection in VHR remote sensing images. IEEE Geoscience and Remote Sensing Letters, 16(2): 310-314
  20. 20.
    Wang J Q, Chen K, Yang S, Loy C C and Lin D H. 2019b. Region proposal by guided anchoring//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE: 2960-2969
  21. 21.
    Xiao B, Wu H P and Wei Y C. 2018. Simple baselines for human pose estimation and tracking//15th European Conference on Computer Vision. Munich: Springer: 472-487
  22. 22.
    Yang X, Yan J C, Feng Z M and He T. 2021. R3Det: refined single-stage detector with feature refinement for rotating object. Proceedings of the AAAI Conference on Artificial Intelligence, 35(4): 3163-3171
  23. 23.
    Zhou X Y, Wang D Q and Krähenbühl P. 2019. Objects as points. arXiv: 1904.07850
  24. 24.
    Zhu X Z, Cheng D Z, Zhang Z, Lin S and Dai J F. 2019. An empirical study of spatial attention mechanisms in deep networks//2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE: 6687-6696

Leer el texto completo

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website