Super-resolution reconstruction of hypertemporal remote sensing images based on self-attention

  • role: First author第一作者
  • Affiliation:

    China Qian Xuesen Space Technology Laboratory, China Academy of Space Technology, Beijing 100094, China

    College of Software, Henan University, Kaifeng 475004, China

  • Email:631719950@qq.com
  • Introduction:E-mail 631719950@qq.com
TANG Xiaotian12,  
  • role: Corresponding author通信作者
  • Affiliation:

    China Qian Xuesen Space Technology Laboratory, China Academy of Space Technology, Beijing 100094, China

    International Institute for Earth System Science, Nanjing University, Nanjing 210023, China

  • Email:yangxue@qxslab.cn
  • Introduction:E-mail yangxue@qxslab.cn
YANG Xue13*,  
  • Affiliation:

    China Qian Xuesen Space Technology Laboratory, China Academy of Space Technology, Beijing 100094, China

LI Feng1,  
  • Affiliation:

    College of Software, Henan University, Kaifeng 475004, China

MA Jun2,  
  • Affiliation:

    Department of Electronic Engineering, Tsinghua University, Beijing 100084, China

LIANG Liang4

resumen

Video satellite hypertemporal data have the characteristics of high temporal resolution, while the single-frame super-resolution reconstruction algorithm can only use the information of the image frame itself, and the reconstruction effect is limited. Therefore, how to utilize fully and effectively the rich spatiotemporal information in hypertemporal data in the super-resolution reconstruction of video satellite images is an issue of interest.Aiming at the characteristics of hypertemporal data, this study proposes a self-attention-based super-resolution reconstruction model of hypertemporal remote sensing images. The model can mine high-frequency information from low-resolution images through an end-to-end network. High-resolution images are recovered from multiple frames of low-resolution images. First, the hypertemporal sequence frames are divided into multiple time groups according to the frame rate, and the spatial and temporal information under different time groups are extracted by using the characteristics of 3D convolution to model space and time simultaneously. It pays attention to the calculation range and completes high dynamic mapping, extracts rich spatial detail information while realizing hypertemporal sequence frame modeling, and finally fuses the features of multiple time groups and completes the reconstruction through subpixel convolution to improve the resolution. The advantage of the proposed algorithm is that the multitime group feature fusion method can extract the spatiotemporal information of sequence frames in multiple time dimensions and fully mine the rich spatiotemporal-related information in the hypertemporal data; the improved self-attention block can be used without registration. It can complete the modeling of sequence frames and improve the extraction ability of detailed spatial information.Experiments on the GF-4 dataset show that the subjective visual effect and objective evaluation index of the algorithm are better than those of the comparison algorithm. When the GF-4 dataset is reconstructed twice, the PSNR value is improved by more than 2.49 dB compared with the bicubic interpolation algorithm, and it still has a strong reconstruction performance when the reconstruction is four times, which is a huge improvement compared with the bicubic interpolation algorithm. Experimental results show that the method has good super-resolution reconstruction effect, which is beneficial to the application of hypertemporal data in various fields.The self-attention-based super-resolution model of hypertemporal remote sensing images proposed in this study fully extracts the spatiotemporal information in hypertemporal data by dividing multiple time groups and calculating the attention features in each time group. The combination of multitemporal group feature fusion and self-attention enables the modeling of overphase sequence frames while ensuring the ability to extract detailed information. Comparative experiments on the GF-4 dataset show that the algorithm in this study is superior to the compared algorithm in terms of objective evaluation indicators and subjective visual effects, verifying the effectiveness and advancement of the algorithm in super-resolution reconstruction of hypertemporal data and the algorithm’s improved reconstruction performance. However, the proposed algorithm must be optimized in terms of calculation time. In a follow-up research, the algorithm structure will be optimized (e.g., changing the residual structure of the network) to reduce the calculation time while ensuring accuracy.

palabra clave

remote sensing;hyper-temporal data;super-resolution reconstruction;deep learning;fusion feature of multiple time groups;wide self-attention;GF-4

References

  1. 1.
    Arefin M R, Michalski V, St-Charles P L, Kalaitzis A, Kim S, Kahou S E and Bengio Y. 2020. Multi-image super-resolution for remote sensing using deep recurrent networks//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Seattle: IEEE: 816-825
  2. 2.
    Caballero J, Ledig C, Aitken A, Acosta A, Totz J, Wang Z, and Shi W. 2017. Real -time video super -resolution with spatio -temporal networks and motion compensation//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4778-4787[]
  3. 3.
    Dai Z H, Liu H X, Le Q V and Tan M X. 2021. CoAtNet: marrying convolution and attention for all data sizes//Proceedings of the Advances in Neural Information Processing Systems 34. [s.l.]: NeurIPS: 3965-3977
  4. 4.
    Dong C, Loy C C, He K M and Tang X O. 2016. Image super-resolution using deep convolutional networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(2): 295-307
  5. 5.
    Dong X Y, Sun X, Jia X P, Xi Z H, Gao L R and Zhang B. 2021. Remote sensing image super-resolution using novel dense-sampling networks. IEEE Transactions on Geoscience and Remote Sensing, 59(2): 1618-1633
  6. 6.
    Fan Z H, Dan T T, Yu H H, Liu B Y and Cai H M. 2020. Single fundus image super-resolution via cascaded channel-wise attention network//42nd Annual International Conference of the IEEE Engineering in Medicine and Biology Society. Montreal: IEEE: 1984-1987
  7. 7.
    Haut J M, Fernandez-Beltran R, Paoletti M E, Plaza J and Plaza A. 2019. Remote sensing image superresolution using deep residual channel attention. IEEE Transactions on Geoscience and Remote Sensing, 57(11): 9277-9289
  8. 8.
    He Z and He D. 2020. Deep learning-based super-resolution for GF-4 satellite imagery. Journal of Remote Sensing (in Chinese), 24(12): 1500-1510
  9. 9.
    Jiang K, Wang Z Y, Yi P, Wang G C, Lu T and Jiang J J. 2019. Edge-enhanced GAN for remote sensing image superresolution. IEEE Transactions on Geoscience and Remote Sensing, 57(8): 5799-5812
  10. 10.
    Kingma D P and Ba J. 2015. Adam: A method for stochastic optimization//3rd International Conference on Learning Representations. San Diego: ICLR
  11. 11.
    Lai W S, Huang J B, Ahuja N and Yang M H. 2019. Fast and accurate image super-resolution with deep Laplacian pyramid networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11): 2599-2613
  12. 12.
    Li F, Yang X, Lu X T, Xin L, Lu M and Zhang N. 2021. A new hyper-temporal imaging mode for spaceborne CMOS cameras. National Remote Sensing Bulletin, 25(1): 514-525
  13. 13.
    Lim B, Son S, Kim H, Nah S and Lee K M. 2017. Enhanced deep residual networks for single image super-resolution//2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops. Honolulu: IEEE: 1132-1140
  14. 14.
    Molini A B, Valsesia D, Fracastoro G and Magli E. 2020. DeepSUM: deep neural network for super-resolution of unregistered multitemporal images. IEEE Transactions on Geoscience and Remote Sensing, 58(5): 3644-3656
  15. 15.
    Nie J, Deng L, Hao X L, Liu M and He Y. 2018. Application of GF-4 satellite in drought remote sensing monitoring: a case study of Southeastern Inner Mongolia. Journal of Remote Sensing (in Chinese), 22(3): 400-407
  16. 16.
    Shi W Z, Caballero J, Huszar F, Totz J, Aitken A P, Bishop R, Rueckert D and Wang Z H. 2016. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network//2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 1874-1883
  17. 17.
    Srinivas A, Lin T Y, Parmar N, Shlens J, Abbeel P and Vaswani A. 2021. Bottleneck transformers for visual recognition//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE: 16514-16524
  18. 18.
    Tian Y P, Zhang Y L, Fu Y and Xu C L. 2020. TDAN: temporally-deformable alignment network for video super-resolution//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE: 3357-3366
  19. 19.
    Tran D, Bourdev L, Fergus R, Torresani L and Paluri M. 2015. Learning spatiotemporal features with 3D convolutional networks//2015 IEEE International Conference on Computer Vision. Santiago: IEEE: 4489-4497
  20. 20.
    Vaswani A, Shazzer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L and Polosukhin I. 2017. Attention is all you need//Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook: Curran Associates Inc.: 5998-6008
  21. 21.
    Wang L G, Guo Y L, Liu L, Lin Z P, Deng X P and An W. 2020. Deep video super-resolution using HR optical flow estimation. IEEE Transactions on Image Processing, 29: 4323-4336
  22. 22.
    Xu L N and He L X. 2017. GF-4 images super resolution reconstruction based on POCS. Acta Geodaetica et Cartographica Sinica, 46(8): 1026-1033
  23. 23.
    Yang X, Li F, Lu M, Xin L, Lu X T and Zhang N. 2022. New super-resolution reconstruction method based on Mixed Sparse Representations. National Remote Sensing Bulletin, 26(8): 1685-1697
  24. 24.
    Yang Y and Newsam S. 2010. Bag-of-visual-words and spatial extensions for land-use classification//Proceedings of the 18th SIGSPATIAL International Conference on Advances in Geographic Information Systems. San Jose: ACM: 270-279
  25. 25.
    Ying X Y, Wang L G, Wang Y Q, Sheng W D, An W and Guo Y L. 2020. Deformable 3D convolution for video super-resolution. IEEE Signal Processing Letters, 27: 1500-1504
  26. 26.
    Zhang Y L, Li K P, Li K, Wang L C, Zhong B N and Fu Y. 2018. Image super-resolution using very deep residual channel attention networks//2018 15th European Conference on Computer Vision. Munich: Springer: 294-310

Leer el texto completo

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website