Feature reassembly and self-attention for oriented object detection in remote sensing images

  • role: First author第一作者
  • Affiliation:

    School of Electronic Information, Northwestern Polytechnical University, Xi'an, 710072, China

  • Email:minlingtong@nwpu.edu.cn
  • Introduction:闵令通,研究方向为人工智能、智能感知和模式识别。E-mail: minlingtong@nwpu.edu.cn
MIN Lingtong1,  
  • Affiliation:

    School of Electronic Information, Northwestern Polytechnical University, Xi'an, 710072, China

FAN Ziman1,  
  • Affiliation:

    School of Automation, Northwestern Polytechnical University, Xi'an, 710072, China

XIE Xingxing2,  
  • role: Corresponding author通信作者
  • Affiliation:

    School of Electronic Information, Northwestern Polytechnical University, Xi'an, 710072, China

  • Email:lvqinyi@nwpu.edu.cn
  • Introduction:吕勤毅,研究方向为人工智能、雷达信息处理和智能感知。E-mail: lvqinyi@nwpu.edu.cn
LYU Qinyi1*

ملخص

Oriented object detection in remote sensing images is an exceptionally challenging task that has elicited widespread attention. With the rapid advancement of deep learning, neural networks based on convolutional neural networks and self-attention networks (e.g., Transformers) have achieved remarkable progress in oriented object detection. However, the focus on boundary and salient feature information in oriented objects in remote sensing images is lacking. Specifically, extracting boundary information for objects with varying orientations is difficult, and the global dependency of salient features is sparse. To address these issues, we propose a method of small-object detection in remote sensing images on the basis of feature reassembly and self-attention. This method consists of a regression branch that incorporates spatial channel reassembly and a self-attention classification branch. The regression branch reassembles spatial information along the channel dimension and emphasizes boundary-sensitive information to achieve accurate localization of bounding boxes. The classification branch leverages self-attention with positional information to capture fundamentally discriminative object features, thus enhancing global feature dependencies for precise classification. Extensive experiments demonstrate the effectiveness and robustness of the proposed model and showcase its excellent performance on publicly available datasets, such as DOTA, HRSC2016, and SODA-A.

مفهوم

remote sensing image;small object detection;detection head;feature reorganization;transformer

References

  1. 1.
    Azimi S M, Vig E, Bahmanyar R, Körner M and Reinartz P. 2018. Towards multi-class object detection in unconstrained remote sensing imagery//Proceedings of the 14th Asian Conference on Computer Vision. Perth: Springer: 150-165
  2. 2.
    Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A and Zagoruyko S. 2020. End-to-end object detection with transformers//Proceedings of the 16th European Conference on Computer Vision. Glasgow: Springer: 213-229
  3. 3.
    Chen Z M, Chen K A, Lin W Y, See J, Yu H, Ke Y and Yang Y. 2020. PIoU loss: towards accurate oriented object detection in complex environments//Proceedings of the 16th European Conference on Computer Vision. Glasgow: Springer: 195-211
  4. 4.
    Cheng G, Li Q Y, Wang G X, Xie X X, Min L T and Han J W. 2023a. SFRNet: fine-grained oriented object recognition via separate feature refinement. IEEE Transactions on Geoscience and Remote Sensing, 61: 5610510
  5. 5.
    Cheng G, Yao Y Q, Li S Y, Li K, Xie X X, Wang J B, Yao X W and Han J W. 2022. Dual-aligned oriented detector. IEEE Transactions on Geoscience and Remote Sensing, 60: 5618111
  6. 6.
    Cheng G, Yuan X, Yao X W, Yan K B, Zeng Q H, Xie X X and Han J W. 2023b. Towards large-scale small object detection: survey and benchmarks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(11): 13467-13488
  7. 7.
    Ding J, Xue N, Long Y, Xia G S and Lu Q K. 2019. Learning RoI transformer for oriented object detection in aerial images//Proceedings of 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE
  8. 8.
    Everingham M, Van Gool L, Williams C K I, Winn J and Zisserman A. 2010. The pascal visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2): 303-338
  9. 9.
    Fu K, Chang Z H, Zhang Y and Sun X. 2021. Point-based estimator for arbitrary-oriented object detection in aerial images. IEEE Transactions on Geoscience and Remote Sensing, 59(5): 4370-4387
  10. 10.
    Han J M, Ding J, Li J and Xia G S. 2022. Align deep features for oriented object detection. IEEE Transactions on Geoscience and Remote Sensing, 60: 5602511
  11. 11.
    He K M, Gkioxari G, Dollár P and Girshick R. 2017. Mask R-CNN//Proceedings of 2017 IEEE International Conference on Computer Vision. Venice: IEEE
  12. 12.
    Jiang Y Q, Tan Z Y, Wang J Y, Sun X Y, Lin M and Li H. 2022. GiraffeDet: a heavy-neck paradigm for object detection. arXiv preprint arXiv: 2202.04256
  13. 13.
    Jiang Y Y, Zhu X Y, Wang X B, Yang S L, Li W, Wang H, Fu P and Luo Z B. 2017. R2CNN: rotational region CNN for orientation robust scene text detection. arXiv preprint arXiv: 1706.09579
  14. 14.
    Li C Z, Xu C Y, Cui Z, Wang D, Zhang T and Yang J. 2019. Feature-attentioned object detection in remote sensing imagery//Proceedings of 2019 IEEE International Conference on Image Processing. Taipei, China: IEEE
  15. 15.
    Li H G, Yu R N and Ding W R. 2021. Research development of small object traching based on deep learning. Acta Aeronauticaet Astronautica Sinica, 42(7): 024691
  16. 16.
    Li W T, Chen Y J, Hu K X and Zhu J K. 2022. Oriented RepPoints for aerial object detection//Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE
  17. 17.
    Lin T Y, Goyal P, Girshick R, He K M and Dollár P. 2017. Focal loss for dense object detection//Proceedings of 2017 IEEE International Conference on Computer Vision. Venice: IEEE
  18. 18.
    Lin T Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P and Zitnick C L. 2014. Microsoft COCO: common objects in context//Proceedings of the 13th European Conference on Computer Vision. Zurich: Springer: 740-755
  19. 19.
    Liu Z K, Wang H Z, Weng L B and Yang Y P. 2016. Ship rotated bounding box space for ship extraction from high-resolution optical satellite images with complex backgrounds. IEEE Geoscience and Remote Sensing Letters, 13(8): 1074-1078
  20. 20.
    Ma J Q, Shao W Y, Ye H, Wang L, Wang H, Zheng Y B and Xue X Y. 2018. Arbitrary-oriented scene text detection via rotation proposals. IEEE Transactions on Multimedia, 20(11): 3111-3122
  21. 21.
    Min L T, Fan Z M, Lv Q Y, Reda M, Shen L H and Wang B L. 2023. YOLO-DCTI: small object detection in remote sensing base on contextual transformer enhancement. Remote Sensing, 15(16): 3970
  22. 22.
    Ming Q, Zhou Z Q, Miao L J, Zhang H W and Li L H. 2021. Dynamic anchor learning for arbitrary-oriented object detection//Proceedings of the 35th AAAI Conference on Artificial Intelligence. [s.l.]: AAAI Press
  23. 23.
    Nie G T and Huang H. 2021. A survey of object detection in optical remote sensing images. Acta Automatica Sinica, 47(8): 1749-1768
  24. 24.
    Nie G T and Huang H. 2023. Multi-oriented object detection in aerial images with double horizontal rectangles. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4): 4932-4944
  25. 25.
    Pan X J, Ren Y Q, Sheng K K, Dong W M, Yuan H L, Guo X W, Ma C Y and Xu C S. 2020. Dynamic refinement network for oriented and densely packed object detection//Proceedings of 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE
  26. 26.
    Qian W, Yang X, Peng S L, Yan J C and Guo Y. 2021. Learning modulated loss for rotated object detection//Proceedings of the 35th AAAI Conference on Artificial Intelligence. [s.l.]: AAAI Press
  27. 27.
    Song G L, Liu Y and Wang X G. 2020. Revisiting the sibling head in object detector//Proceedings of 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE
  28. 28.
    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones K, Gomez A N, Kaiser Ł and Polosukhin I. 2017. Attention is all you need//Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach: Curran Associates Inc.: 6000-6010
  29. 29.
    Wang J W, Ding J, Guo H W, Cheng W S, Pan T and Yang W. 2019. Mask OBB: a semantic attention-based mask oriented bounding box representation for multi-category object detection in aerial images. Remote Sensing, 11(24): 2930
  30. 30.
    Wang J W, Yang W, Li H C, Zhang H J and Xia G S. 2021. Learning center probability map for detecting objects in aerial images. IEEE Transactions on Geoscience and Remote Sensing, 59(5): 4307-4323
  31. 31.
    Xia G S, Bai X, Ding J, Zhu Z, Belongie S, Luo J B, Datcu M, Pelillo M and Zhang L P. 2018. DOTA: a large-scale dataset for object detection in aerial images//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE
  32. 32.
    Xie X X, Cheng G, Wang J B, Yao X W and Han J W. 2021. Oriented R-CNN for object detection//Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE
  33. 33.
    Xu Y C, Fu M T, Wang Q M, Wang Y K, Chen K, Xia G S and Bai X. 2021. Gliding vertex on the horizontal bounding box for multi-oriented object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(4): 1452-1459
  34. 34.
    Yang X, Yan J C, Feng Z M and He T. 2021. R3Det: refined single-stage detector with feature refinement for rotating object//Proceedings of the AAAI 35th Conference on Artificial Intelligence. [s.l.]: AAAI Press
  35. 35.
    Yang X, Yang J R, Yan J C, Zhang Y, Zhang T F, Guo Z, Sun X and Fu K. 2019. SCRDet: towards more robust detection for small, cluttered and rotated objects//Proceedings of 2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE
  36. 36.
    Zhang G J, Lu S J and Zhang W. 2019. CAD-Net: a context-aware detection network for objects in remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 57(12): 10015-10024
  37. 37.
    Zhang G J, Luo Z P, Tian Z C, Zhang J Y, Zhang X Q and Lu S J. 2023. Towards efficient use of multi-scale features in transformer-based object detectors//Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE

قراءة النص الكامل

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website