Semantic understanding of geo-objects’ relationship in high resolution remote sensing image driven by dual LSTM

  • role: First author第一作者
  • Affiliation:

    School of Geosciences and Info-physics, Central South University, Changsha 410083, China

  • Email:cj2011@csu.edu.cn
  • Introduction:1980E-mailcj2011@csu.edu.cn
CHEN Jie,  
  • Affiliation:

    School of Geosciences and Info-physics, Central South University, Changsha 410083, China

DAI Xinyi,  
  • Affiliation:

    School of Geosciences and Info-physics, Central South University, Changsha 410083, China

ZHOU Xing,  
  • Affiliation:

    School of Geosciences and Info-physics, Central South University, Changsha 410083, China

SUN Geng,  
  • role: Corresponding author通信作者
  • Affiliation:

    School of Geosciences and Info-physics, Central South University, Changsha 410083, China

  • Email:dengmin@csu.edu.cn
  • Introduction:1974E-maildengmin@csu.edu.cn
DENG Min*

résumé

Geo-objects in High-resolution Remote Sensing Images (HRSIs) have clear category attributes and rich semantic information. With the support of artificial intelligence technology, the spatial relationship can be automatically recognized by a computer. At present, the semantic understanding of HRSIs mainly relies on an image caption model to generate sentences based on the global features. However, coarse-grained features can easily cause the category attribute of the object to be mispredicted during the sentence generation process. In fact, taking the geo-object as the basic unit of semantic understanding is more in line with people’s habit of cognizing geographic space. To obtain more accurate sentences, this study constructs an Object-based Geo-spatial Relation Image Understanding Dataset (OGRIUD) and proposes a dual LSTM-driven semantic understanding method.The proposed dataset is based on the object, and the sentence description includes the category and location information of the ground object, which make up the deficiency of the target category and the location information in the semantic understanding of the current remote sensing field. The proposed method uses the object detection model to identify salient objects in the image and uses the object features as input in the language model to alleviate the problem of incorrectly predicted categories in the description. Furthermore, to use HRSI scene information, we fuse the global and regional features and use dual LSTM to predict the attention distribution of each geo-object.We compare the global feature-based approach with the object feature based approach proposed in this paper. Quantitative analysis results show that the proposed method exhibits increased exact matching accuracy, from 53.5% of the original to 62.33%. The visual analysis results show that the proposed method, and the generated spatial relation description statements are also more abundant.This method enables the language model to focus on objects with actual semantics, and the matching degree between the generated description statement and the remote sensing image content is also improved. The correspondence between the visual object and description improves the interpretability of remote sensing image understanding.

mots-clés

high resolution remote sensing image;ground objects;spatial relationships;semantic understanding;image caption

References

  1. 1.
    Anderson P, He X D, Buehler C, Teney D, Johnson M, Gould S and Zhang L. 2018. Bottom-up and top-down attention for image captioning and visual question answering//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE: 6077-6086
  2. 2.
    Chen J, Han Y R, Wan L, Zhou X and Deng M. 2019. Geospatial relation captioning for high-spatial-resolution images by using an attention-based neural network. International Journal of Remote Sensing, 40(16): 6482-6498
  3. 3.
    Cui W, Zhang D Y, He X, Yao M, Wang Z W, Hao Y J, Li J, Wu W J, Cui W Q and Huang J J. 2019. Multi-scale remote sensing semantic analysis based on a global perspective. ISPRS International Journal of Geo-Information, 8(9): 417
  4. 4.
    Karpathy A and Li F F. 2015. Deep visual-semantic alignments for generating image descriptions//Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE: 3128-3137
  5. 5.
    Lin C Y. 2004. Looking for a few good metrics: ROUGE and its evaluation//Proceedings of the 4th NTCIR Workshop. Tokyo: National Institute of Informatics: 1-8
  6. 6.
    Lu X Q, Wang B Q, Zheng X T and Li X L. 2018. Exploring models and data for remote sensing image caption generation. IEEE Transactions on Geoscience and Remote Sensing, 56(4): 2183-2195
  7. 7.
    Papineni K, Roukos S, Ward T and Zhu W J. 2002. Bleu: a method for automatic evaluation of machine translation//Proceedings of the 40th Annual Meeting on Association for Computational Linguistics. Philadelphia: Association for Computational Linguistics: 311-318
  8. 8.
    Qu B, Li X L, Tao D C and Lu X Q. 2016. Deep semantic understanding of high resolution remote sensing image//2016 International Conference on Computer, Information and Telecommunication Systems. Kunming: IEEE: 1-5
  9. 9.
    Shi Z W and Zou Z X. 2017. Can a machine generate humanlike language descriptions for a remote sensing image?. IEEE Transactions on Geoscience and Remote Sensing, 55(6): 3623-3634
  10. 10.
    Vedantam R, Zitnick C L and Parikh D. 2015. CIDEr: consensus-based image description evaluation//Proceddings of 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE: 4566-4575
  11. 11.
    Xu K, Ba J L, Kiros R, Cho K, Courville A, Salakhutdinov R, Zemel R S and Bengio Y. 2015. Show, attend and tell: neural image caption generation with visual attention//Proceedings of the 32nd International Conference on International Conference on Machine Learning. Lille: JMLR.org: 2048-2057
  12. 12.
    Yang Y and Newsam S. 2010. Bag-of-visual-words and spatial extensions for land-use classification//Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems. San Jose: Association for Computing Machinery: 270-279
  13. 13.
    You Q Z, Jin H L, Wang Z W, Fang C and Luo J B. 2016. Image captioning with semantic attention//Proceddings of 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 4651-4659
  14. 14.
    Yuan Z H, Li X L and Wang Q. 2020. Exploring multi-level attention and semantic relationship for remote sensing image captioning. IEEE Access, 8: 2608-2620
  15. 15.
    Zhang F, Du B and Zhang L P. 2015. Saliency-guided unsupervised feature learning for scene classification. IEEE Transactions on Geoscience and Remote Sensing, 53(4): 2175-2184
  16. 16.
    Zhang Z Y, Diao W H, Zhang W K, Yan M L, Gao X and Sun X. 2019a. LAM: remote sensing image captioning with label-attention mechanism. Remote Sensing, 11(20): 2349
  17. 17.
    Zhang Z Y, Zhang W K, Diao W H, Yan M L, Gao X and Sun X. 2019b. VAA: visual aligning attention model for remote sensing image captioning. IEEE Access, 7: 137355-137364

Lire l'article complet

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website