- 1.
Anderson P, He X D, Buehler C, Teney D, Johnson M, Gould S and Zhang L. 2018. Bottom-up and top-down attention for image captioning and visual question answering//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE: 6077-6086
- 2.
Chen J, Han Y R, Wan L, Zhou X and Deng M. 2019. Geospatial relation captioning for high-spatial-resolution images by using an attention-based neural network. International Journal of Remote Sensing, 40(16): 6482-6498
- 3.
Cui W, Zhang D Y, He X, Yao M, Wang Z W, Hao Y J, Li J, Wu W J, Cui W Q and Huang J J. 2019. Multi-scale remote sensing semantic analysis based on a global perspective. ISPRS International Journal of Geo-Information, 8(9): 417
- 4.
Karpathy A and Li F F. 2015. Deep visual-semantic alignments for generating image descriptions//Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE: 3128-3137
- 5.
Lin C Y. 2004. Looking for a few good metrics: ROUGE and its evaluation//Proceedings of the 4th NTCIR Workshop. Tokyo: National Institute of Informatics: 1-8
- 6.
Lu X Q, Wang B Q, Zheng X T and Li X L. 2018. Exploring models and data for remote sensing image caption generation. IEEE Transactions on Geoscience and Remote Sensing, 56(4): 2183-2195
- 7.
Papineni K, Roukos S, Ward T and Zhu W J. 2002. Bleu: a method for automatic evaluation of machine translation//Proceedings of the 40th Annual Meeting on Association for Computational Linguistics. Philadelphia: Association for Computational Linguistics: 311-318
- 8.
Qu B, Li X L, Tao D C and Lu X Q. 2016. Deep semantic understanding of high resolution remote sensing image//2016 International Conference on Computer, Information and Telecommunication Systems. Kunming: IEEE: 1-5
- 9.
Shi Z W and Zou Z X. 2017. Can a machine generate humanlike language descriptions for a remote sensing image?. IEEE Transactions on Geoscience and Remote Sensing, 55(6): 3623-3634
- 10.
Vedantam R, Zitnick C L and Parikh D. 2015. CIDEr: consensus-based image description evaluation//Proceddings of 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE: 4566-4575
- 11.
Xu K, Ba J L, Kiros R, Cho K, Courville A, Salakhutdinov R, Zemel R S and Bengio Y. 2015. Show, attend and tell: neural image caption generation with visual attention//Proceedings of the 32nd International Conference on International Conference on Machine Learning. Lille: JMLR.org: 2048-2057
- 12.
Yang Y and Newsam S. 2010. Bag-of-visual-words and spatial extensions for land-use classification//Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems. San Jose: Association for Computing Machinery: 270-279
- 13.
You Q Z, Jin H L, Wang Z W, Fang C and Luo J B. 2016. Image captioning with semantic attention//Proceddings of 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 4651-4659
- 14.
Yuan Z H, Li X L and Wang Q. 2020. Exploring multi-level attention and semantic relationship for remote sensing image captioning. IEEE Access, 8: 2608-2620
- 15.
Zhang F, Du B and Zhang L P. 2015. Saliency-guided unsupervised feature learning for scene classification. IEEE Transactions on Geoscience and Remote Sensing, 53(4): 2175-2184
- 16.
Zhang Z Y, Diao W H, Zhang W K, Yan M L, Gao X and Sun X. 2019a. LAM: remote sensing image captioning with label-attention mechanism. Remote Sensing, 11(20): 2349
- 17.
Zhang Z Y, Zhang W K, Diao W H, Yan M L, Gao X and Sun X. 2019b. VAA: visual aligning attention model for remote sensing image captioning. IEEE Access, 7: 137355-137364