References
- 1.Akiva P, Purri M and Leotta M. 2022. Self-supervised material and texture representation learning for remote sensing tasks//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE: 8193-8205
- 2.Ayush K, Uzkent B, Meng C L, Tanmay K, Burke M, Lobell D and Ermon S. 2021. Geography-aware self-supervised learning//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE: 10161-10170
- 3.Bi K F, Xie L X, Zhang H H, Chen X, Gu X T and Tian Q. 2022. Pangu-weather: a 3D high-resolution model for fast and accurate global weather forecast. arXiv preprint arXiv: 2211.02556
- 4.Bourcier J, Floquet T, Dashyan G, Ceillier T, Alahari K and Chanussot J. 2022. Self-supervised pretraining on satellite imagery: a case study on label-efficient vehicle detection. arXiv preprint arXiv: 2210.11815
- 5.Cha K, Seo J and Lee T. 2023. A billion-scale foundation model for remote sensing images. arXiv preprint arXiv: 2304.05215
- 6.Chen K, Han T, Gong J C, Bai L, Ling F H, Luo J J, Chen X, Ma L M, Zhang T N, Su R, Ci Y Z, Li B, Yang X K and Ouyang W L. 2023. FengWu: pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv: 2304.02948
- 7.Chen T, Kornblith S, Norouzi M and Hinton G. 2020a. A simple framework for contrastive learning of visual representations//Proceedings of the 37th International Conference on Machine Learning. Virtual: JMLR.org: 1597-1607
- 8.Chen T, Kornblith S, Swersky K, Norouzi M and Hinton G. 2020b. Big self-supervised models are strong semi-supervised learners//Proceedings of the 34th International Conference on neural Information Processing Systems. Vancouver: Curran Associates Inc.: 22243-22255
- 9.Chen X L, Fan H Q, Girshick R and He K M. 2020c. Improved baselines with momentum contrastive learning. arXiv preprint arXiv: 2003.04297
- 10.Chen X, Xie S, He K. An empirical study of training self-supervised vision transformers[C]//Proceedings of the IEEE/CVF international conference on computer vision. 2021: 9640-9649.
- 11.Cong Y Z, Khanna S, Meng C L, Liu P, Rozi E, He Y T, Burke M, Lobell D B and Ermon S. 2022. SatMAE: pre-training transformers for temporal and multi-spectral satellite imagery//Proceedings of the 36th International Conference on Neural Information Processing Systems. New Orleans: Curran Associates Inc.: 197-211
- 12.Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X H, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J and Houlsby N. 2020. An image is worth 16×16 words: transformers for image recognition at scale. arXiv preprint arXiv: 2010.11929
- 13.Gómez C, White J C and Wulder M A. 2016. Optical remotely sensed time series data for land cover classification: a review. ISPRS Journal of Photogrammetry and Remote Sensing, 116: 55-72
- 14.Guibas J, Mardani M, Li Z Y, Tao A, Anandkumar A and Catanzaro B. 2021. Adaptive fourier neural operators: efficient token mixers for transformers. arXiv preprint arXiv: 2111.13587
- 15.He K M, Chen X L, Xie S N, Li Y H, Dollár P and Girshick R. 2022. Masked autoencoders are scalable vision learners//Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE: 15979-15988
- 16.He K M, Fan H Q, Wu Y X, Xie S N and Girshick R. 2020. Momentum contrast for unsupervised visual representation learning//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE: 9726-9735
- 17.Heidler K, Mou L C, Hu D, Jin P, Li G Y, Gan C, Wen J R and Zhu X X. 2023. Self-supervised audiovisual representation learning for remote sensing data. International Journal of Applied Earth Observation and Geoinformation, 116: 103130
- 18.Ienco D, Interdonato R, Gaetano R and Minh D H T. 2019. Combining Sentinel-1 and Sentinel-2 Satellite Image Time Series for land cover mapping via a multi-source deep learning architecture. ISPRS Journal of Photogrammetry and Remote Sensing, 158: 11-22
- 19.Jain P, Schoen-Phelan B and Ross R. 2021. Multi-modal self-supervised representation learning for earth observation//2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS. Brussels: IEEE: 3241-3244 [DOI: .
- 20.2021.9553741]
- 21.Jain P, Schoen-Phelan B and Ross R. 2022. Self-supervised learning for invariant representations from multi-spectral and SAR images. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 15: 7797-7808
- 22.Jung H, Oh Y, Jeong S, Lee C and Jeon T. 2022. Contrastive self-supervised learning with smoothed representation for remote sensing. IEEE Geoscience and Remote Sensing Letters, 19: 8010105
- 23.Lam R, Sanchez-Gonzalez A, Willson M, Wirnsberger P, Fortunato M, Alet F, Ravuri S, Ewalds T, Eaton-Rosen Z, Hu W H, Merose A, Hoyer S, Holland G, Vinyals O, Stott J, Pritzel A, Mohamed S and Battaglia P. 2022. GraphCast: learning skillful medium-range global weather forecasting. arXiv preprint arXiv: 2212.12794
- 24.Li W Y, Chen K Y, Chen H and Shi Z W. 2022a. Geographical knowledge-driven representation learning for remote sensing images. IEEE Transactions on Geoscience and Remote Sensing, 60: 5405516
- 25.Li W Y, Chen K Y and Shi Z W. 2022b. Geographical supervision correction for remote sensing representation learning. IEEE Transactions on Geoscience and Remote Sensing, 60: 5411520
- 26.Li Z, Sui Z W, Fu Q Y, Zheng J J and Bu T. 2023. High-resolution remote sensing extraction of urban buildings based on morphological sequences and multi-source a priori information. National Remote Sensing Bulletin, 27(4): 998-1008
- 27.Mai G C, Lao N, He Y T, Song J M and Ermon S. 2023. CSP: self-supervised contrastive spatial pre-training for geospatial-visual representations. //International Conference on Machine Learning. PMLR, 2023: 23498-23515
- 28.Mall U, Hariharan B and Bala K. 2023. Change-aware sampling and contrastive learning for satellite images//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE: 5261-5270
- 29.Mañas O, Lacoste A, Giró-i-Nieto X, Vazquez D and Rodríguez P. 2021. Seasonal contrast: unsupervised pre-training from uncurated remote sensing data//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE: 9394-9403
- 30.Mendieta M, Han B R, Shi X J, Zhu Y and Chen C. 2023. GFM: building geospatial foundation models via continual pretraining. arXiv preprint arXiv: 2302.04476
- 31.Muhtar D, Zhang X L, Xiao P F, Li Z S and Gu F. 2023. CMID: a unified self-supervised learning framework for remote sensing image understanding. IEEE Transactions on Geoscience and Remote Sensing, 61: 5607817
- 32.Pathak J, Subramanian S, Harrington P, Raja S, Chattopadhyay A, Mardani M, Kurth T, Hall D, Li Z Y, Azizzadenesheli K, Hassanzadeh P, Kashinath K and Anandkumar A. 2022. FourCastNet: a global data-driven high-resolution weather model using adaptive fourier neural operators. arXiv preprint arXiv: 2202.11214
- 33.Patnala A, Stadtler S, Schultz M G and Gall J. 2023. Generating views using atmospheric correction for contrastive self-supervised learning of multispectral images. IEEE Geoscience and Remote Sensing Letters, 20: 2502305
- 34.Prexl J and Schmitt M. 2023. Multi-modal multi-objective contrastive learning for Sentinel-1/2 imagery//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Vancouver: IEEE: 2136-2144
- 35.Reed C J, Gupta R, Li S F, Brockman S, Funk C, Clipp B, Keutzer K, Candido S, Uyttendaele M and Darrell T. 2023. Scale-MAE: a scale-aware masked autoencoder for multiscale geospatial representation learning. arXiv preprint arXiv: 2212.14532
- 36.Shi X J, Chen Z R, Wang H, Yeung D Y, Wong W K and Woo W C. 2015. Convolutional LSTM network: a machine learning approach for precipitation nowcasting//Proceedings of the 28th International Conference on Neural Information Processing Systems. Montreal: MIT Press: 802-810
- 37.Stewart A J, Lehmann N, Corley I A, Wang Y, Chang Y C, Braham N A A, Sehgal S, Robinson C and Banerjee A. 2023. SSL4EO-L: datasets and foundation models for landsat imagery. arXiv preprint arXiv: 2306.09424
- 38.Stojnić V and Risojević V. 2021. Self-supervised learning of remote sensing scene representations using contrastive multiview coding//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Nashville: IEEE: 1182-1191
- 39.Sun X, Wang P J, Lu W X, Zhu Z C, Lu X N, He Q B, Li J X, Rong X E, Yang Z J, Chang H, He Q L, Yang G, Wang R P, Lu J W and Fu K. 2023. RingMo: a remote sensing foundation model with masked image modeling. IEEE Transactions on Geoscience and Remote Sensing, 61: 5612822
- 40.Tao C, Qi J, Zhang G, Zhu Q, Lu W P and Li H F. 2023. TOV: the original vision model for optical remote sensing image understanding via self-supervised learning. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 16: 4916-4930
- 41.Tian Y L, Krishnan D and Isola P. 2020. Contrastive multiview coding//16th European Conference on Computer Vision. Glasgow: Springer: 776-794
- 42.Tian Z Z, Zhang H W, Wang K, Liu S Q, Zou Q J, Zhao Z and Chen Y B. 2023. Application of an improved CenterNet in remote sensing images object detection. National Remote Sensing Bulletin, 27(12): 2706-2715
- 43.Tseng G, Cartuyvels R, Zvonkov I, Purohit M, Rolnick D and Kerner H. 2024. Lightweight, pre-trained transformers for remote sensing timeseries. arXiv preprint arXiv: 2304.14065
- 44.Vincenzi S, Porrello A, Buzzega P, Cipriano M, Fronte P, Cuccu R, Ippoliti C, Conte A and Calderara S. 2021. The color out of space: learning self-supervised representations for earth observation imagery//2020 25th International Conference on Pattern Recognition (ICPR). Milan: IEEE: 3034-3041
- 45.Wang D, Zhang Q M, Xu Y F, Zhang J, Du B, Tao D C and Zhang L P. 2022a. Advancing plain vision transformer toward remote sensing foundation model. IEEE Transactions on Geoscience and Remote Sensing, 61: 5607315
- 46.Wang W, Li X J and Wang X. ADC-CPANet:A Remote Sensing Image Classification Method Based on Local-Global Feature Fusion. National Remote Sensing Bulletin,
- 47.Wang Y B, Wu H X, Zhang J J, Gao Z F, Wang J M, Yu P S and Long M S. 2022b. PredRNN: a recurrent neural network for spatiotemporal predictive learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2): 2208-2225
- 48.Wanyan X Y, Seneviratne S, Shen S C and Kirley M. 2023. DINO-MC: self-supervised contrastive learning for remote sensing imagery with multi-sized local crops. arXiv preprint arXiv: 2303.06670
- 49.Ying C X, Cai T L, Luo S J, Zheng S X, Ke G L, He D, Shen Y M and Liu T Y. 2021. Do transformers really perform bad for graph representation?. arXiv:2106.05234
- 50.Yuan Y and Lin L. 2021. Self-supervised pretraining of transformers for satellite image time series classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14: 474-487
- 51.Yuan Y, Lin L, Liu Q S, Hang R L and Zhou Z G. 2022. SITS-Former: a pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classification. International Journal of Applied Earth Observation and Geoinformation, 106: 102651
- 52.Zheng X C, Kellenberger B, Gong R, Hajnsek I and Tuia D. 2021. Self-supervised pretraining and controlled augmentation improve rare wildlife recognition in UAV images//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops. Montreal: IEEE: 732-741


