Revue des recherches interdisciplinaires entre intelligence artificielle et science de la télédétection : état des lieux et perspectives

  • role: First author第一作者
  • Affiliation:

    State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430072, China

  • Email:weihe1990@whu.edu.cn
  • Introduction:/E-mail weihe1990@whu.edu.cn
HE Wei1,  
  • Affiliation:

    State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430072, China

WANG Zijie1,  
  • role: Corresponding author通信作者
  • Affiliation:

    College of Oceanography and Space Informatics, China University of Petroleum (East China), Qingdao 266580, China

  • Email:genyunsun@163.com
  • Introduction:E-mailgenyunsun@163.com
SUN Genyun2*,  
  • Affiliation:

    School of Automation, Northwestern Polytechnical University, Xi'an 710072, China

CHENG Gong3,  
  • Affiliation:

    School of Artificial Intelligence, Xidian University, Xi'an 710068, China

TANG Xu4,  
  • Affiliation:

    State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430072, China

WU Chen1,  
  • Affiliation:

    School of Resources and Civil Engineering, Northeastern University, Shenyang 110819, China

HE Liming5,  
  • Affiliation:

    School of Earth and Space Sciences, Peking University, Beijing 100871, China

REN Huazhong6,  
  • Affiliation:

    School of Remote Sensing & Geomatics Engineering, Nanjing University of Information Science & Technology, Nanjing 210044, China

HU Ting7,  
  • Affiliation:

    College of Information and Communication Engineering, Harbin Engineering University, Harbin 266404, China

FENG Shou8,  
  • Affiliation:

    Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China

NIE Sheng9,  
  • Affiliation:

    Advanced Institute for Satellite Application, Beijing Normal University, Beijing 100088, China

WU Shangrong10,  
  • Affiliation:

    College of Oceanography and Space Informatics, China University of Petroleum (East China), Qingdao 266580, China

GAO Han2,  
  • Affiliation:

    School of Artificial Intelligence, Xidian University, Xi'an 710068, China

FENG Jie4,  
  • Affiliation:

    School of Computer Science and Technology, Nanjing University of Information Science & Technology, Nanjing 210044, China

HANG Renlong11,  
  • Affiliation:

    School of Artificial Intelligence, Anhui University, Hefei 230601, China

DING Yun12,  
  • Affiliation:

    School of Information Engineering, North China University of Water Resources and Electric Power, Zhengzhou 450046, China

ZHANG Rui13,  
  • Affiliation:

    Faculty of Geosciences and Environmental Engineering, Southwest Jiaotong University, Chengdu 611756, China

YE Yuanxin14,  
  • Affiliation:

    Faculty of Geosciences and Environmental Engineering, Southwest Jiaotong University, Chengdu 611756, China

MA Xianping14,  
  • Affiliation:

    Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China

ZHAO Dan9,  
  • Affiliation:

    College of Geomatics, Shandong University of Science and Technology, Qingdao 266590, China

LI Zhenhai15,  
  • Affiliation:

    Institute of Digital China (Fujian), Fuzhou University, Fuzhou 350108, China

SU Hua16,  
  • Affiliation:

    College of Architecture and Urban Planning, Shenzhen University, Shenzhen 518060, China

XU Nan17,  
  • Affiliation:

    School of Geography and Geomatics Engineering, Suzhou University of Science and Technology, Suzhou 215009, China

CHEN Chao18,  
  • Affiliation:

    State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430072, China

MA Ailong1,  
  • Affiliation:

    School of Geography and Information Engineering, China University of Geosciences (Wuhan), Wuhan 430074, China

ZHU Qiqi19,  
  • Affiliation:

    Advanced Institute for Satellite Application, Beijing Normal University, Beijing 100088, China

YAN Kai10,  
  • Affiliation:

    Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences, Changchun 130102, China

JIA Mingming20,  
  • Affiliation:

    Department of Geography, The University of Hong Kong, Hong Kong 999077, China

ZHANG Hongsheng21,  
  • Affiliation:

    Faculty of Geography, Yunnan Normal University, Kunming 650050, China

LUO Yi22

résumé

Le développement rapide de l'intelligence artificielle pousse la science de la télédétection à passer d'un paradigme « observation dominante » à un paradigme « cognition intelligente ». Face aux défis posés par l'expansion rapide de l'échelle d'observation caractérisée par la multi-source hétérogène et les caractéristiques de haute dimension, les méthodes traditionnelles d'interprétation peinent à répondre aux exigences pratiques en termes d'efficacité, de précision et d'évolutivité. Les technologies d'intelligence artificielle représentées par l'apprentissage profond et les grands modèles offrent un nouveau support théorique et des voies techniques pour l'extraction automatique de caractéristiques, la fusion multimodale et la découverte de connaissances en profondeur dans les systèmes terrestres complexes. Ces dernières années, la fusion des données de télédétection avec intelligence artificielle s'est approfondie dans plusieurs types de données d'observation (optique haute résolution, hyperspectrale, SAR, LiDAR, etc.) et diverses tâches intelligentes (classification, détection, segmentation, détection de changement et raisonnement avec grands modèles), démontrant un potentiel pour remodeler les modes de cognition et renforcer l'intelligence décisionnelle dans des scénarios d'application clés tels que la géologie, l'écologie, l'agriculture, l'urbanisme et la surveillance des catastrophes. Cependant, la recherche actuelle présente encore des problèmes tels qu'un couplage insuffisant entre les mécanismes d'observation et la représentation des modèles, une capacité limitée de généralisation interrégionale et intermodale, ainsi que des défis en termes d'interprétabilité et de fiabilité des systèmes intelligents. Sur cette base, cet article examine systématiquement les derniers progrès de l'intelligence artificielle au service de la science de la télédétection selon trois dimensions : technologies d'observation, méthodes intelligentes et applications typiques, résume leur évolution et caractéristiques communes, explore les défis clés liés aux systèmes terrestres dynamiques complexes à l'échelle mondiale, et envisage la construction d'un nouveau système théorique intelligent pour la télédétection de prochaine génération, généralisable, interprétable et durable.

mots-clés

science de la télédétection;intelligence artificielle;apprentissage profond;applications interdisciplinaires;big data de télédétection;interprétation intelligente

1 Introduction

As a science that conducts wide area observations of the Earth's surface through non-contact methods, remote sensing provides support for human understanding and comprehension of Earth system processes. With the rapid development and widespread deployment of satellite imagery, aerial photography, and ground sensors, the scale of remote sensing data is growing exponentially, driving Earth observation into the "big data era". However, while large-scale, multi-source, and heterogeneous remote sensing data greatly expands Earth observation capabilities, it also poses unprecedented challenges to traditional data acquisition, processing, analysis, and interpretation models, and urgently requires the construction of a new generation of intelligent theoretical and technological systems to support it.
The rapid development of artificial intelligence technology in the field of computer vision has brought a new paradigm for automated processing of high-dimensional data and complex pattern recognition. Its ability for autonomous learning and representation has further propelled a crucial leap from observation to cognition. These advantages make artificial intelligence an ideal technology for solving high-dimensional, large-scale, and complex nonlinear problems in current remote sensing science. In recent years, an increasing number of studies have introduced cutting-edge artificial intelligence methods into the field of remote sensing.Fig. 1The trend in the number of research papers on the intersection of artificial intelligence and remote sensing science over the past decade is presented, covering Remote Sensing of Environment、ISPRS Journal of Photogrammetry and Remote Sensing、IEEE Transactions on Geoscience and Remote Sensing Waiting for top international journals, as well as important domestic journals such as the Journal of Remote Sensing and the Journal of Surveying and Mapping. It can be clearly seen that since 2021, the number of publications and citations of related papers has significantly increased, indicating that remote sensing science research driven by artificial intelligence has become a hot topic and frontier in the field.Table 1Further summarized the main research tasks of artificial intelligence in remote sensing science, covering three major levels: remote sensing observation technology (high-resolution remote sensing, hyperspectral remote sensing, radar remote sensing, LiDAR, thermal infrared, and low altitude remote sensing); Intelligent analysis methods (remote sensing classification, object detection, semantic segmentation, change detection, large models and embodied intelligence); Typical application scenarios (geological remote sensing, ecological remote sensing, wetland remote sensing, agricultural remote sensing, ocean/coastal remote sensing, disaster remote sensing, and urban remote sensing). This systematic development pattern indicates that artificial intelligence is profoundly reshaping the theoretical framework and technological system of remote sensing science.
figure
figure

Fig. 1 Statistics of the number of articles related to remote sensing science and artificial intelligence published in important domestic and foreign journals in the past decade (The statistics presented herein were obtained from Web of Science)

Table 1 Representative tasks in the field of remote sensing science with artificial intelligence technology

研究层面任务种类典型方法数据来源
遥感观测技术高分辨率遥感CNNGF-2、GL-1、GeoEye-1、WorldView-2等
高光谱遥感3D CNNGF-5、EO-1、Zhuhai-1、航空平台等
雷达遥感随机森林、CNNGF-3、Sentinel-1、TerraSAR-X等
激光雷达3D CNNICESat-2、GEDI、GF-7、勾芒号等
热红外遥感CNNMODIS、HJ-2A、SDGSAT-1等
低空遥感迁移学习无人机数据等
智能分析方法遥感分类CNN、TransformerGF-2、Sentinel-2、Google Earth images等
目标检测YOLO、DETRGF-2、JL-1、Google Earth images,等
语义分割UNetJL-1、GF-2、Google Earth images、航空平台等
变化检测CNN、RNNGF-2、WorldView-2/3、Sentinel-2等
大模型对比学习、预训练—微调多模态遥感图像、自然语言指令等
具身智能强化学习点云地图、轨迹路径、环境状态变化等
典型应用场景地质遥感随机森林、CNNSentinel-2、Landsat、GF-5、Sentinel-1、SRTM等
生态遥感DNNMODIS NDVI/EVI、HJ-1A/1B CCD、GIMMS等
湿地遥感LSTM、CNNSentinel、Landsat、MODIS、GDEM、SRTM等
农业遥感LSTM、机理模型MODIS、Sentinel-2、HJ-1A/1B CCD等
海洋/海岸遥感ANN、CNNMODIS-Aqua/Terra、GHRSST、HY-2等
灾害遥感CNN、TransformerSentinel-1/2、WorldView-2/3/4、Himawari-8、OSM等
城市遥感CNNSentinel-2、Landsat、Sentinel-1、GF系列等
Although the integration of artificial intelligence and remote sensing science has shown significant cross innovation trends, existing research is mostly focused on specific algorithms or single application scenarios, and there is still a lack of systematic theoretical frameworks and panoramic analysis. To compensate for this deficiency, it is necessary to analyze and construct a comprehensive research report covering multi-source observations, intelligent methods, and typical applications. Therefore, this article systematically reviews the latest progress and development trends of artificial intelligence empowering remote sensing science, from observation techniques, algorithm methods to application scenarios. asFig. 2Firstly, based on the perspective of remote sensing observation technology, analyze the key issues in information extraction and feature analysis of different observation methods, such as high-resolution remote sensing, hyperspectral remote sensing, radar remote sensing, LiDAR, etc., and explore new ideas for artificial intelligence in modeling complex physical processes and improving interpretation accuracy. Secondly, from the perspective of artificial intelligence technology, this paper examines the adaptability and transformative potential of cutting-edge methods such as deep learning, large models, and embodied intelligence in remote sensing science, revealing their advantages in representation learning, knowledge transfer, and cross modal perception. Furthermore, at the application level, this paper systematically elaborates on the paradigm shift driven by artificial intelligence in typical scenarios such as geology, ecology, oceanography, cities, and disasters, demonstrating its significant value in deepening understanding of the Earth system and supporting intelligent decision-making. Finally, this article delves into the core challenges that constrain the deep integration of artificial intelligence and remote sensing, and looks forward to possible future research directions and breakthrough paths.
figure

Fig. 2 Overview framework diagram (From remote sensing observation to artificial intelligence technology and then to cross-application)

Cross development of artificial intelligence and remote sensing science

From the perspective of the ontology logic of remote sensing observation, the improvement of Earth system observation capability fundamentally depends on the depth and complementarity of different observation methods in spatial, spectral, and temporal dimensions. Currently, multi-source observation systems represented by high-resolution optical remote sensing, hyperspectral imaging, radar remote sensing, and LiDAR are capable of achieving complementary support in areas such as fine surface structure, material composition, dynamic changes, and three-dimensional geometry. These sensors have significant differences in imaging mechanism, spatial resolution, spectral response, and spatiotemporal coverage, leading to issues such as multimodal information fusion, feature analysis, and uncertainty propagation, becoming a key bottleneck for high-precision remote sensing interpretation(Guo和Liang,2024). In the face of nonlinear, multi-scale, and highly coupled Earth observation problems, traditional analysis methods that rely on physical models and artificial experience are no longer sufficient to fully characterize their complex mechanisms. The introduction of artificial intelligence provides an important opportunity for remote sensing science to shift from experience driven to intelligent driven. By leveraging the hierarchical feature representation and cross modal correlation modeling of deep networks, artificial intelligence can demonstrate potential beyond traditional methods in complex physical process representation, feature pattern abstraction, and information fusion optimization. This reshapes the basic paradigm of remote sensing science at the three cognitive levels of "how to observe", "what to observe", and "how to understand"(Fig. 3). Representative datasets of six core directions in remote sensing science and artificial intelligence technology, such asTable 2As shown.
figure

Fig. 3 Artificial intelligence empowers the paradigm evolution of multi-source remote sensing from observation to cognition

Table 2 Representative datasets in the field of remote sensing science with artificial intelligence technology

遥感类型典型数据集数据规模主要来源核心观测特征主要应用场景
高分辨率遥感DOTA v2.011268图像航空/商业卫星亚米级空间分辨率目标检测、精细结构解析
SpaceNet系列数万km²Maxar/Planet建筑与道路几何语义/实例分割、灾害评估
FAIR1M15000图像多源卫星精细目标外观高分遥感目标检测
LoveDA5987图像Google Earth城市场景结构跨区域语义分割
高光谱遥感Indian Pines145×145×200AVIRIS高维光谱信息物质识别、光谱解混
Pavia University610×340×103ROSIS城市光谱差异精细分类
雷达遥感SEN12MS180662图像对Sentinel-1/2结构与介电特性光学—SAR融合
OSCD24区域Sentinel-2/SAR时序变化城市变化检测
激光雷达遥感ISPRS Vaihingen33图像航空LiDAR高精度三维几何3D 语义分割
ISPRS Potsdam38图像LiDAR/光学高程与形态三维结构建模
Semantic3D数十亿点地面 LiDAR超密集点云点云语义理解
热红外遥感KAIST Multispectral95328图像对航空平台温度与热辐射热目标检测
FLIR ADAS v29711图像车载TIR热异常感知夜间目标识别
Landsat TIR全球长期NASA地表温度热异常与环境监测
低空遥感UAVid30视频无人机超高分辨率城市场景理解
UAVDT10 h无人机低空动态视角目标检测与跟踪

2.1 High resolution remote sensing

High resolution remote sensing, with its excellent ability to capture ground details, has become one of the core technologies for modern Earth observation and environmental monitoring. Compared with medium and low resolution remote sensing, its spatial resolution usually reaches sub meter level or even higher, which can provide fine surface structure information and provide strong data support for applications such as urban modeling, precision agriculture, infrastructure monitoring, and disaster assessment. Typical data sources include high-resolution optical satellites such as the WorldView series GF-2/6)、 Airborne LiDAR and unmanned aerial vehicle imaging platforms are capable of acquiring high spatial, multispectral, and even hyperspectral data at different scales and perspectives. In terms of data dimension features, high-resolution remote sensing can also provide multi band and multi polarization information in the spectral dimension, combined with its high timeliness and repetitive coverage ability, achieving fine observation of surface dynamic changes.
High resolution remote sensing in urban planning(Amini等,2022)Land use/cover monitoring(Robinson等,2019)Smart Transportation(Agarwal等,2022)Environmental monitoring and disaster response(Ahmad,2024)Such fields have shown extensive potential for application. For example, the acquisition of high-precision building and road information in cities strongly supports the construction of digital twin cities; Fine monitoring of agricultural plots and assessment of crop growth status provide scientific decision-making basis for precision agriculture; In disaster monitoring, high-resolution remote sensing can quickly locate damaged areas and provide timely response for rescue deployment and loss assessment.
However, the in-depth application of high-resolution remote sensing still faces several development bottlenecks. Firstly, the massive amount of data places high demands on storage, transmission, and computing resources; Secondly, the imaging process is susceptible to interference from factors such as noise, cloud cover, lighting changes, and atmospheric scattering, which reduces the stability and interpretability of the data; In addition, the high cost of manual annotation of high-resolution data leads to a shortage of large-scale high-quality training samples, which restricts the development of fine target recognition and multi-source fusion analysis capabilities(Jiang等,2022).
In response to the above issues, artificial intelligence technology provides a new solution for the processing and analysis of high-resolution remote sensing data. Through a multimodal learning framework, high-resolution optical images, LiDAR point clouds, and radar data can achieve cross modal information fusion and collaborative modeling, effectively enhancing target detection, change monitoring, and 3D reconstruction capabilities in complex scenes(Hong等,2023Luo等,2024). In addition, deep learning has demonstrated unique advantages in real-time processing, cross scale object recognition, and automated analysis of large-scale data, providing key support for promoting the leap of high-resolution remote sensing from "data acquisition" to "intelligent perception"(Li等,2025a).

2.2 Hyperspectral remote sensing

Hyperspectral remote sensing, as an advanced Earth observation technology, relies on imaging spectrometers to obtain reflection and radiation information of ground objects in hundreds of continuous and narrow spectral bands(Fu等,2023Hu等,2025Zhao等,2021). This technology breaks through the limitations of traditional multispectral and achieves feature expression of "graph integration" in the form of an "image cube", allowing each pixel to have a fine and continuous spectral curve. This feature enhances the ability to finely characterize land features and identify features that are difficult to distinguish in multispectral remote sensing, such as mineral absorption valleys and vegetation red edge effects(Petersson等,2016Yu等,2022). Based on the above advantages, hyperspectral remote sensing has been widely applied in major fields such as precision agriculture, environmental monitoring, geological exploration, and disaster assessment(Hosseinpour-Zarnaq等,2025Pande和Moharir,2023).
However, the in-depth application of hyperspectral remote sensing technology still faces several key challenges. Firstly, its high-dimensional characteristics result in a massive amount of data, which places higher demands on data storage, transmission, and computing resources. Secondly, during the imaging process, factors such as lighting, terrain, atmospheric effects, and sensor noise collectively constrain the spectral stability of remote sensing data. More importantly, the scarcity of real terrain annotation samples, coupled with inherent complexities such as "same spectrum foreign objects" and "same object but different spectra" in the data, severely restricts the automation and high-precision interpretation of hyperspectral information(Kherimiche等,2024).
In response to the limitations and insufficient generalization ability of traditional methods in processing high-dimensional and nonlinear hyperspectral data, deep learning technology has injected new development momentum into this field. In response to the problem of difficult correlation between spectra and space in hyperspectral image classification, LKSSAN(Sun等,2023a)The effective modeling of long-range 3D feature dependencies using large kernel attention and convolutional feedforward structure significantly improves classification performance while maintaining the 3D structure. With its powerful automatic feature extraction capability, deep learning models can achieve end-to-end extraction of spectral spatial joint features, effectively capturing complex nonlinear relationships, thereby significantly improving the accuracy and robustness of land cover interpretation. Furthermore, by introducing multimodal learning mechanisms, deep learning models can integrate hyperspectral data with cross modal data such as LiDAR and multispectral, enhancing information complementarity and interpretation robustness(Hu等,2023aZhao 等,2023b). In recent years, with the continuous accumulation of large-scale hyperspectral datasets and the rise of large-scale modeling techniques such as HyperSigma(Wang等,2025a)The HyperFree(Li等,2025c))Hyperspectral remote sensing is gradually achieving a paradigm shift from data acquisition to intelligent interpretation(He等,2022).

2.3 Radar remote sensing

Radar remote sensing uses microwave bands for detection, which can obtain various characteristic parameters such as backscattering coefficient, polarization decomposition characteristics, and radar index. Due to the complex nonlinear relationship between microwave characteristics and parameters, machine learning models rely on their high flexibility and ability to handle high-dimensional nonlinear problems(Ali等,2015)Has become the mainstream method in vegetation parameter SAR inversion research(Kumar等,2018)For example, random forest(Ahmadian等,2019Kumar等,2018)Support Vector Regression(Xie等,2021)Gaussian Process Regression(Xu等,2022a)Wait. Research has shown that multi feature training is the foundation for the effective implementation of machine learning algorithms, and combining the collaborative input of radar and optical features can further improve the accuracy of parameter inversion(Bahrami等,2021Xu等,2022a).
More importantly, the rapid development of artificial intelligence is driving a paradigm shift in polarimetric SAR data processing from traditional manual interpretation to intelligent automation. Deep learning technology significantly improves the accuracy and efficiency of crop classification and forest monitoring by delving into the complex correlations between polarization features and land cover attributes. Deep learning models represented by Convolutional Neural Networks (CNN)(Kim,2014)It can directly learn scattering physical properties from raw pixels without relying on manual feature design, demonstrating significant advantages in processing large-scale complex polarization data. With the continuous emergence of architectures such as deep Boltzmann machines, graph convolutional networks, and lightweight networks based on attention mechanisms (such as DeepLabV3+)(Shi等,2025)The development of complex valued neural networks also provides a new technological path for fully exploring complex valued polarization information(Alkhatib等,2025).
However, existing inversion models that rely mainly on experience or semi experience have weakened the in-depth exploration of radar features and the intrinsic scattering mechanism between vegetation canopies as the research paradigm gradually shifts from mechanistic models to data-driven models. Therefore, developing data-driven models that integrate microwave scattering mechanisms has become an important topic. In classification tasks, difficulties in sample annotation, scarcity of publicly available datasets, and high-dimensional model parameters constrain the generalization of methods. To alleviate the dependence on annotated data, recent research has focused on techniques such as semi supervised learning, small sample learning, transfer learning, and domain adaptation. In addition, lightweight models and physical neural networks guided by scattering mechanisms are becoming important research directions for improving model generalization ability, accuracy, and interpretability.

2.4 Lidar remote sensing

Lidar, as a typical active remote sensing technology, can efficiently obtain the three-dimensional geometric shape and echo intensity information of ground objects by emitting laser pulses and measuring their round-trip time(Wang等,2024b). This technology supports deployment on multiple platforms including ground, vehicle, airborne, and satellite, meeting the multi-scale observation needs from local to global. Typical missions such as ICESat-2, GEDI, Gaofen-7, and Goumang have achieved global vertical structure measurements; The airborne and unmanned aerial vehicle LiDAR systems provide high-density point cloud data at local and regional scales, supporting multi-level applications from macro terrain mapping to fine 3D reconstruction.
Lidar data has significant multidimensional information representation capabilities: in spatial dimensions, its elevation and positioning accuracy can reach centimeter to meter levels; In terms of time dimension, support multi period observations to characterize surface dynamic changes; In terms of signal dimension, it covers intensity, multiple echoes, full waveform, and even polarization parameters. In addition, some multi/hyperspectral LiDAR systems can also provide spectral information in both visible and near-infrared bands simultaneously(Gao等,2025). These multidimensional features collectively endow LiDAR with unique advantages in structured, quantitative response, and multi-scale surface cognition.
With its high precision and 3D representation capabilities, LiDAR has played an important role in fields such as forestry, terrain mapping, urban modeling, and disaster monitoring(Wang等,2024bDong和Chen,2017). In ecological monitoring, LiDAR can accurately extract three-dimensional structural parameters of the canopy, providing key data support for biomass and carbon storage estimation; In terrain surveying, even in areas with dense vegetation coverage, high-precision digital elevation models (DEM) and digital surface models (DSM) can still be generated; In the field of smart cities, LIDAR point clouds provide precise spatial basis for the construction of digital twins; In terms of disaster response, differential analysis based on multi period LiDAR data can effectively identify landslides, ground subsidence, and surface deformation. In addition, this technology has shown continuous potential for application in coastal and wetland monitoring, infrastructure inspection, and other scenarios.
However, the further development of LiDAR technology still faces several key bottlenecks(杨必胜和董震,2019Wang等,2024b). Firstly, massive point cloud data poses significant pressure on storage, transmission, and computing resources; Secondly, the common noise interference, redundant information, and irregular structure in the data significantly increase the complexity of subsequent interpretation and 3D modeling; In addition, the lack of large-scale high-quality annotated samples also constrains the accuracy and generalization ability of deep learning algorithms. Traditional methods based on artificial features and empirical rules lack robustness in complex scenarios, making it difficult to achieve highly robust automated interpretation.
With the rapid development of artificial intelligence technology, the deep integration of LiDAR and deep learning has become an important direction for intelligent remote sensing. Deep neural networks can automatically extract multi-level spatial geometric and semantic features, significantly improving the accuracy and robustness of object detection, semantic segmentation, and parameter inversion(Dong和Chen,2017). Meanwhile, the end-to-end learning framework driven by artificial intelligence, combined with GPU parallel computing capabilities, effectively achieves a significant improvement in point cloud processing efficiency. Looking ahead to the future, integrating LiDAR with multimodal data such as multi/hyperspectral and radar to achieve cross modal feature collaboration and knowledge transfer will become a key direction for further enhancing LiDAR interpretation capabilities and expanding its intelligent application boundaries(Magruder等,2024).

2.5 Thermal infrared remote sensing

With the continuous advancement of thermal infrared sensor technology, satellites, drones, and ground platforms can easily obtain thermal infrared remote sensing data. The research on target recognition based on thermal infrared data, combined with artificial intelligence technology, has gradually attracted widespread attention. At present, there are multiple publicly available ground platform thermal infrared target detection datasets, such as FLIR ADAS Dataset v1.3 and V2.1, LLVIP, OSU Thermal Pedestrian Database, M3FD, and CVC-14(Davis和Sharma,2007González等,2016Jia等,2021), as well as ground view video sequence datasets for thermal infrared target tracking tasks, such as BU-TIV Dataset, PTB-TIR, PDT-ATV, Terravic Motion IR, and LSOTB-TIR(Liu等,2020Portmann等,2014). However, thermal infrared target detection datasets collected based on unmanned aerial vehicle platforms are relatively scarce. Existing datasets include the BIRDSAI dataset built by teams such as Harvard University, the VEDAI dataset from the University of Caen in France, and the DroneVehicle dataset released by Tianjin University(Bondi等,2020Jiang等,2025Razakarivony和Jurie,2016Sun等,2022).
At present, research on thermal infrared remote sensing target detection mainly focuses on ground perspective thermal infrared images, which are mainly applied in fields such as autonomous driving, unmanned driving, and medical diagnosis. Due to the similarity in imaging angle between ground platform thermal infrared images and large-scale optical natural image datasets, researchers typically fine tune pre trained models of large-scale optical datasets through deep learning(Bochkovskiy等,2020Zhou等,2021)To reduce the demand for data volume and adapt to new thermal infrared target detection tasks. The transfer learning based thermal infrared remote sensing target detection method effectively reduces the training time and computational resource requirements, while improving the model's generalization ability. However, aerial thermal infrared remote sensing target detection faces challenges in complex scenarios. Firstly, due to the complex imaging conditions, the target features are often not obvious, and in some cases, the background information is even more prominent than the target, resulting in a large number of missed and false detections. Secondly, the difference in target size, especially the size variation under sensors of different resolutions or the same resolution, poses extremely high requirements for the scale adaptability of detection algorithms. Developing detection algorithms that balance both small-scale and large-scale targets remains a major challenge in current research. In addition, there are significant differences in signal-to-noise ratio and spatial resolution between images obtained by different thermal infrared sensors(Jiang等,2024).
Compared with thermal infrared data from ground and unmanned aerial vehicle platforms, the spatial resolution of spaceborne thermal infrared remote sensing is relatively low, mostly at the kilometer level. The highest publicly available resolution is currently at the ten meter level. Therefore, research on target detection in spaceborne thermal infrared remote sensing mainly focuses on cloud detection tasks, with common methods including threshold segmentation and techniques based on feature extraction and classification(Liu等,2011Zhu和Woodcock,2012). Thermal infrared remote sensing target detection also plays an important role in military applications, especially in areas such as camouflage target detection, ship recognition, infrared guidance, and mine detection, which are of great significance. Disguised object detection based on deep learning has become a research hotspot in the field of computer vision, which not only enhances natural image object detection technology, but also promotes the intelligent process of military monitoring applications(李建东 等,2024李召良 等,2025史彩娟 等,2022).

2.6 Low altitude remote sensing

Low altitude remote sensing relies on drones and light aviation platforms to obtain surface information in airspace below 1000 meters above the ground. Compared to satellite remote sensing, it has significant advantages: high resolution can reach centimeter or even millimeter level(李林源 等,2025)Capable of finely capturing land features; Capable of high-frequency acquisition and rapid response, able to achieve repeated coverage in a short period of time; The platform has strong flexibility and supports the installation of multiple sensors to meet the application needs of various fields such as agricultural monitoring and urban planning(吴志峰 等,2025); At the same time, it has the characteristics of low cost and easy operation, and can quickly obtain customized data; It can still maintain high reliability under complex meteorological conditions, overcoming the limitation of satellite remote sensing being easily affected by weather.
Currently, low altitude remote sensing datasets are developing towards a systematic approach of "task scene modality". Task dimensions cover object detection and tracking (such as VisDrone)(Zhu等,2022a)The DroneSwarms(Cao等,2025))Semantic segmentation (such as UAVid)(Lyu等,2020)The SkyScenes(Khose等,2025))Instance counting (such as DroneCrowd)(Wen等,2021)The AnimalDrone(Zhu等,2021))And autonomous navigation (such as AerialVLN)(Liu等,2023c)The CityNav(Lee等,2025b))Waiting for intelligent perception tasks. The scenario dimension covers urban transportation (such as AU-AIR)(Bozcan和Kayacan,2020)The DRIFT(Lee等,2025a))Smart agriculture (such as PDT)(Zhou等,2025)The WeedsGalore(Celikkan等,2025))Maritime supervision (such as SeaDronesSee-v2)(Varga等,2022)The KOLOMVERSE(Nanda等,2024))And security emergencies (such as Anti UAV)(Jiang等,2023)The CST Anti-UAV(Xie等,2025))Waiting for multiple fields. The modal dimension extends from a single visible light modality to multimodal data fusion such as thermal infrared and LiDAR point clouds (such as HIT-UAV)(Suo等,2023)The MARS-LVIG(Nanda等,2024)The UAVScenes(Wang等,2025e)), Significantly improved the robustness and cross domain adaptability of data, providing a solid foundation for the intelligent application of low altitude remote sensing.
With the continuous development of 3D reconstruction technology, low altitude remote sensing is gradually evolving towards the construction of 3D datasets. 3D reconstruction effectively breaks through the expression limitations of 2D images and can construct refined 3D environmental models, providing more accurate spatial data support for urban planning, cultural heritage protection, and other fields. Combining virtual reality and augmented reality technology, 3D data will further promote the implementation of application scenarios such as intelligent transportation and virtual navigation, enhancing its practical value in digital twins and emergency response.
Although low altitude remote sensing has obvious advantages in spatial resolution and flexibility, it still faces challenges such as large data volume, scarce annotated samples, and significant sensor noise. Massive data requires enormous storage and computation, especially in the process of multimodal and 3D data fusion. The lack of consistent labeling and high-quality samples constrains the generalization ability of the model. In addition, factors such as changes in lighting and meteorological conditions can also affect the stability of remote sensing data, further increasing the complexity of data interpretation.
The introduction of deep learning has brought new breakthroughs to low altitude remote sensing. Its automatic feature extraction capability supports end-to-end joint learning of spectral spatial features, significantly improving the accuracy and robustness of land cover interpretation. Especially with the integration of multimodal learning mechanisms, intelligent collaboration between low altitude remote sensing, LiDAR, optical imaging and other multi-dimensional data has been achieved, enhancing the integrity and interpretation accuracy of information. With the continuous development of large-scale datasets and big model technologies (such as HyperSigma)(Wang等,2025a)The HyperFree(Li等,2025c))Low altitude remote sensing is accelerating towards a new stage of intelligent interpretation, and promoting continuous technological progress in areas such as unmanned system collaborative perception and disaster response.

3 Interdisciplinary Technologies of Artificial Intelligence and Remote Sensing Science

From the cognitive logic of intelligent analysis, the knowledge production of remote sensing science is undergoing a fundamental transformation from experience driven to intelligent driven. With the continuous improvement of remote sensing observation resolution and data diversity, traditional interpretation methods that rely on manual experience and physical modeling have become difficult to achieve efficient and high-precision information extraction when dealing with complex Earth system processes that are multi-scale, strongly nonlinear, and cross modal. The introduction of artificial intelligence has opened up a new cognitive path for remote sensing information processing. With the hierarchical representation and pattern abstraction capabilities of deep networks, remote sensing data can transition from static feature interpretation to dynamic semantic understanding, achieving intelligent collaboration throughout the entire process from land cover classification, target recognition, semantic segmentation to change detection. Furthermore, the new generation of technological paradigms represented by large models and embodied intelligence is expanding the cognitive boundaries of remote sensing science. The large model constructs a unified perception framework through universal representation and cross domain transfer, endowing the model with strong generalization performance across regions and sensors; Embodied intelligence integrates real-time 3D and digital twin technology, promoting remote sensing from passive observation to active cognition and decision-making reasoning. Therefore, the deep integration of artificial intelligence and remote sensing science is not only an upgrade in methodology, but also a paradigm reconstruction of the observation, cognition, and understanding of the Earth system. This section will focus on the six core directions of classification, detection, segmentation, change detection, large models, and embodied intelligence, systematically sorting out the technological context and cutting-edge progress of this paradigm shift (such asFig. 4As shown).
figure

Fig. 4 Related artificial intelligence techniques used in remote sensing science

3.1 Remote sensing classification technology

Remote sensing image classification technology identifies the category and spatial distribution of objects based on their electromagnetic radiation characteristics in the image, in order to quickly grasp the regional environmental conditions(Maulik和Chakraborty,2017). The accuracy of remote sensing image classification is influenced by factors such as detection channels, spectral response of ground objects, atmospheric propagation characteristics, and sensor performance. Artificial intelligence technology has broken through the assumption dependence of traditional methods on data distribution, and utilizes convolution operations in deep neural networks to automatically extract remote sensing image features, effectively suppressing the aforementioned interference factors and providing a new technological approach for remote sensing image classification.
The widely used artificial intelligence algorithms in this field currently include CNN, graph neural networks, and Transformer networks(Veličković等,2018)Mamba Network(Gu和Dao,2024)And the large model(Liu等,2023b)Wait. CNN, with its local receptive field and weight sharing mechanism, can effectively extract advanced features of remote sensing data, and continuously optimize feature expression ability through error backpropagation, constructing a nonlinear mapping from features to categories with good generalization performance. GCN uses message passing mechanism to smooth the feature space and alleviate the impact of spectral variation on classification by constructing a graph structure between labeled and unlabeled samples. The Transformer network relies on self attention mechanism to capture the contextual relationships between samples in images, enhancing the discriminative features of ground features. The Mamba network is based on structured state space sequence modeling, introducing a selection mechanism to achieve efficient filtering of input related information. Combined with hardware aware algorithms, it significantly improves inference speed and flexibility. Remote sensing large-scale models are usually pre trained on a large-scale classification task and used as a foundation. After freezing their parameters, only the top-level classifier is fine tuned, demonstrating excellent transfer and generalization abilities in various downstream tasks.
Although remote sensing classification technology has developed rapidly, it still faces many challenges. Firstly, remote sensing data has strong professionalism and closedness, resulting in a lack of annotated data, which restricts the training and application of data-driven models. Secondly, existing artificial intelligence models suffer from issues such as insufficient reliability and weak interpretability of reasoning, which limit their decision-making credibility in complex scenarios. In addition, the differences in spectral and spatial resolution of different sensors lead to spatio-temporal inconsistency between multi-source data, which further weakens the cross sensor generalization ability of the model.

3.2 Remote sensing target recognition technology

Remote sensing target detection, as a core technology for automatically identifying and locating land targets, has been widely applied in fields such as urban planning, disaster emergency response, environmental monitoring, traffic management, and national defense reconnaissance(Li等,2020Redmon等,2016Ren等,2017Tang等,2025Yang等,2019Yin等,2025).
Early remote sensing object detection mainly relied on manually designed features and shallow learning models, which extracted spectral, texture, shape, and other features and combined them with classifiers to achieve object recognition. This type of method performs well in regular terrain scenes, but has limited detection performance in complex backgrounds, small targets, and coexistence of multiple types of targets. In recent years, the widespread application of deep learning has significantly promoted the intelligent development of remote sensing object detection. CNN has end-to-end feature learning capabilities and can automatically extract multi-layer semantic features from massive data, achieving significant improvements in detection accuracy and robustness. Typical two-stage detection models such as Faster R-CNN(Ren等,2017)PDQR based on dynamic object sequence reconstruction(Yin等,2025)And SCRDet for small targets(Yang等,2019)Significant breakthroughs have been made in terms of accuracy. On the other hand, YOLO(Redmon等,2016)And SSD(Liu等,2016)To represent a one-stage detection algorithm, the object detection task is modeled as a regression problem, which significantly improves inference speed while maintaining high accuracy. YOLO-SS(Tang等,2025)Further optimization is carried out for small targets in remote sensing images, which still has good robustness in complex backgrounds.
The success of deep learning models in remote sensing object detection stems from their highly compatible structure with remote sensing image features. The hierarchical architecture of the roll paper neural network can effectively capture multi-scale spatial patterns of ground objects, and the feature pyramid network significantly enhances the recognition ability of small targets by fusing high-level semantic and low-level detail features; New structures such as Transformer further enhance global context dependent modeling, providing a new technological path for intelligent recognition of land cover in complex scenes.
However, remote sensing object detection still faces multiple challenges. Firstly, complex land cover types and strong background interference lead to feature confusion between the target and background, affecting recognition accuracy; Secondly, high-resolution images generally suffer from weak and high proportion of small target features, resulting in a high rate of missed detections; Thirdly, the scarcity of high-quality annotated samples limits the generalization performance of the model; In addition, significant differences in data distribution across regions and sensors further exacerbate the difficulty of model transferability.
Future research will focus on the following directions: reducing dependence on annotated data through weakly supervised and self supervised learning; Enhance the recognition ability of unknown categories through small sample and zero sample learning; Using cross domain adaptive technology to improve the transfer performance of the model; And based on Transformer and large model architecture, achieve more powerful unified modeling of global features. At the same time, the collaborative integration of target detection, change detection, semantic segmentation and other tasks will promote the paradigm evolution of remote sensing from "passive perception" to "active cognition". Overall, remote sensing object detection is gradually shifting from feature driven to intelligent driven, and its efficiency, generalization, and interpretability will become the core directions of future research.

3.3 Remote sensing segmentation technology

Remote sensing image segmentation aims to achieve pixel level semantic understanding of surface images, parsing complex scenes into semantically meaningful land cover category entities(Li等,2024a;Zhao等,2024). Early methods relied heavily on shallow image features and traditional machine learning algorithms, which had significant shortcomings in accuracy, generalization, and robustness. Fully Convolutional Neural Network (FCN)(Long等,2015)With UNet(Ronneberger等,2015)Since its proposal, deep learning segmentation methods based on encoder decoder structures have rapidly developed and been widely applied in intelligent interpretation of remote sensing images, achieving significant improvements in segmentation accuracy and automation level. In recent years, researchers have made significant progress in multi-scale feature expression, attention mechanisms, and large-scale models.
Typical representatives include ResUNet-a(Diakogiannis等,2020)The HMANet(Niu等,2022)The RSSFormer(Xu等,2023)The UNetFormer(Wang等,2022b)And lightweight remote sensing dedicated frameworks, etc. ResUNet-a(Diakogiannis等,2020)The fusion of residual structure, multi-scale dilated convolution, and pyramid pooling module in the UNet framework effectively enhances the contextual expression ability of features, and improves segmentation accuracy through multi task conditional inference mechanism utilizing semantic associations between tasks(Fig. 5). HMANet(Niu等,2022)On the basis of FCN, class enhancement and region shuffling attention mechanisms are introduced. The former enhances class discrimination by modeling semantic dependencies between pixels, while the latter performs efficient self attention calculations using sparse region representations, thus balancing global modeling and computational efficiency(Fig. 6). RSSFormer(Xu等,2023)Based on the multi-resolution feature flow of HRNet, an adaptive Transformer fusion module is proposed to suppress background noise and highlight foreground saliency during feature aggregation(Fig. 7). UNetFormer(Wang等,2021)The structure combines a lightweight CNN encoder with a global local joint attention Transformer decoder, and captures both window level context and local detail information through a dual branch mechanism(Fig. 8). UrbanSSF(Wang等,2025g)Considering the sequence dependency relationship between features at different stages, a Mamba based decoder is used to model the dynamic interaction between multi-layer features, in order to enhance semantic consistency and spatial structure expression(Fig. 9). In response to the problems of high annotation cost and difficulty in representing heterogeneous features in image segmentation, PAMSNet(Zhao等,2025)A multi-source feature encoding and cross resolution fusion module was proposed, which achieved efficient and accurate segmentation results with only a small number of point annotations.Ma等(2024)Explored SAM (Segment Anything Model)(Kirillov等,2023)In the field of remote sensing, a plug and play lightweight framework is proposed, which significantly improves semantic consistency and boundary accuracy by utilizing SAM generated object and boundary prior information.
figure

Fig. 5 Schematic diagram of the ResUNet-a model

figure

Fig. 6 Schematic diagram of the HMANet model

figure

Fig. 7 Schematic of the adaptive Transformer fusion module in RSSFormer

figure

Fig. 8 Schematic diagram of the UNetFormer model

figure

Fig. 9 Schematic representation of feature state interaction Mamba in UrbanSSF

Overall, with the evolution of deep learning model structures and the rapid development of large models, remote sensing image segmentation research is moving towards the direction of multimodal information fusion, efficient annotation utilization, and strong generalization ability(Li等,2024b)。 Future research urgently needs to break through two bottlenecks: first, reducing reliance on high cost pixel level annotated data to alleviate the performance bottleneck caused by the scarcity of rare category samples; The second is to enhance the robustness and adaptability of the model under cross regional, cross seasonal, and cross sensor conditions, thereby achieving high-precision and transferable segmentation of complex surface scenes, providing solid support for intelligent remote sensing interpretation.

3.4 Remote sensing change detection technology

Remote sensing change detection aims to identify the changing characteristics of surface elements at different time scales, and is an important technical support for environmental monitoring, disaster assessment, and urban dynamic analysis(Li等,2025b). Early methods mainly relied on traditional methods such as image differencing, ratio calculation, principal component analysis, and post classification comparison. However, they were highly sensitive to lighting disturbances, noise interference, and registration errors, which could easily lead to false positives and missed detections. With the development of deep learning, convolutional neural networks (CNN)(Kim,2014)Regarding Transformer Architecture(Vaswani等,2017)Introduced to enhance spatial spectral feature extraction; Recurrent Neural Network (RNN)(Elman,1990)Used to capture time series dependencies; The attention mechanism further enhances the modeling ability of long-range dependencies through adaptive weight allocation, thereby significantly improving the spatial detail preservation and temporal consistency of change detection. In recent years, the rise of large models has broken through the traditional "binary change" and "from to" paradigms, gradually moving change detection towards the intelligent translation stage of cross modal and multi semantic resolution.
The existing research on deep change detection can be summarized into four main technical routes: sample reduction oriented and optimal feature modeling. BCG-Net(Hu等,2023b)Suppressing the propagation of hyperspectral unmixing errors in multi class change detection through the mechanism of "two types of change guidance+time consistency"; STU-SAMI(Zuo等,2025)Using SAM (Segment Anything Model)(Kirillov等,2023)Generate pseudo variation samples, significantly reducing reliance on dual temporal manual labeling; S2C(Ding等,2025)By distilling implicit semantic knowledge from visual models, multimodal representation of changes under zero sample conditions can be achieved; GSTM-SCD (Liu等,2025)Integrating state space and graph structure information to achieve change detection at the dual/multi temporal semantic level. Modeling for spectral uncertainty. QSCDNet(Lv等,2025)Introducing quantum branches to characterize spectral uncertainty in Hilbert space effectively enhances the robustness of hyperspectral change detection. Analysis of continuous temporal changes. Multi-RLD-Net(Li和Wu,2024)By aligning multi temporal images with optical flow, synchronized output of change areas and occurrence times is achieved to achieve frame level fine-grained change localization. Oriented towards semantic interpretability and knowledge enhancement. Prompt-CC(Liu等,2023a)In line with the trend of big models, a two-stage prompt learning framework of "whether to change → why to change" is used to generate natural language change descriptions, achieving unified expression at the pixel level, semantic level, and text level for the first time, promoting the transformation of change detection from "pixel labeling" to "storytelling".
Overall, deep learning technology has significantly promoted the intelligent process of remote sensing change detection, enabling models to have stronger feature expression and semantic understanding capabilities. However, there are still three challenges at present: firstly, the scarcity of high-quality multimodal and multi-phase annotated data limits the generalization ability of the model; Secondly, the spatiotemporal heterogeneity and imaging differences between different sensors lead to insufficient model migration performance; Thirdly, the lack of large models with interpretability and reusability hinders the practical application of tasks in real-world scenarios. Future research can be conducted in two directions: firstly, relying on a large-scale model system to construct a universal change detection large-scale model, achieving adaptive change recognition across regions and disaster types; The second is to combine generative models with semantic description mechanisms to expand the task boundaries of change detection, promoting its evolution from low-level difference detection to high-level semantic parsing and intelligent reasoning.

3.5 Remote Sensing Large Model Technology

In recent years, with the continuous enhancement of remote sensing observation capabilities and the rapid accumulation of multi-source and multi-scale data, remote sensing interpretation technology is undergoing a profound paradigm evolution(An等,2025). Since around 2021, large-scale modeling methods based on massive remote sensing data have gradually matured(Sun等,2023bWang等,2023)This marks a shift in the field from traditional methods relying on manual features and shallow learning to a new stage of intelligent interpretation driven by deep neural networks and big data. Early research focused on single task applications of high-resolution images and multi-source data, such as scene classification and object detection, laying the methodological foundation for the subsequent development of systematic models. In the process of word processing, the remote sensing big model gradually formed four core competency systems: one is the universal representation ability, which can construct stable and comprehensive environmental perception in multi-source, multi-scale, and multi-phase data; The second is cross domain migration capability, which can transfer existing knowledge to new regions, sensors, or tasks, reducing dependence on specific data distributions; The third is the ability for continuous learning, which enables the model to continuously absorb new knowledge under dynamically changing terrain environments and observation conditions, and alleviate catastrophic forgetting problems; The fourth is interpretability ability, which can reveal the inherent logic of model predictions, enhance the credibility and scientific value of interpretation results.
With the rapid development of cross modal learning technologies such as vision and language, multimodal remote sensing modeling has gradually become a research frontier. By integrating multiple heterogeneous data sources such as optical images, SAR images, LiDAR point clouds, and geographic texts, a multimodal remote sensing model is developed(Wang等,2023)Has shown outstanding potential in complex terrain structure recognition, semantic relationship inference, and spatiotemporal dynamic perception. This progress has promoted the transformation of remote sensing interpretation from traditional "task driven" to "scene driven", strengthened the comprehensive modeling of the overall semantics and dynamic evolution process of the geographic environment, and provided important support for achieving high-level understanding of the geographic environment.
Taking the SAM series method as an example, its core innovation lies in using the ViT architecture to build high-capacity encoders and achieving fast segmentation of any target through a "prompt based" mechanism. This design enables SAM to have significant zero sample and few sample generalization capabilities, which can be directly used for fast contour extraction of targets such as buildings, water bodies, and roads in remote sensing scenes. Previous studies have shown that SAM-RS, SAM-CD, and other models that have been adapted or fine tuned exhibit good transfer potential in high-resolution remote sensing segmentation and change detection tasks, especially in scenarios with limited annotated samples or diverse target shapes, which can significantly reduce the cost of manual interaction and training. However, from the perspective of applicable boundaries, native SAM is mainly based on natural image distribution for pre training, and its explicit modeling ability for common scale differences, noise interference (such as SAR speckle noise), and temporal changes in remote sensing images is still limited. It usually requires domain specific feature adaptation, structural constraints, or multimodal information to stably perform in complex remote sensing tasks.
By around 2025, the development of remote sensing large-scale models will enter a systematic and explosive stage. The research focus is gradually shifting towards building a unified architecture to simultaneously support multiple remote sensing tasks such as classification, segmentation, change detection, and object tracking(Hong等,2024Yao等,2023). This type of systematic model enhances adaptability and robustness to complex environments while significantly reducing training costs through shared feature representation and cross task knowledge transfer. At the same time, the widespread adoption of pre training fine-tuning paradigms and self supervised learning methods enables models to efficiently utilize massive amounts of unlabeled data, effectively alleviate the performance limitations caused by scarce labeled samples, and further enhance their generalization and transfer abilities across sensors, regions, and tasks.
With the rise of the concept of general intelligent agents, remote sensing interpretation is gradually evolving towards intelligent systems with comprehensive analysis and decision support capabilities. Previous studies have attempted to integrate visual large-scale models with natural language processing frameworks(Guo等,2024Li等,2025c)Develop a method that can understand task intent, perform multi-step reasoning, and achieve multi-level cognitive analysis of remote sensing data. Such progress indicates that future remote sensing interpretation systems are expected to break through the limitations of single tasks and develop into intelligent decision support tools for complex geographic scenarios, gradually forming a complete technical chain from high-precision information extraction to intelligent decision services, providing solid technical support for smart Earth observation and refined management.

3.6 Remote sensing embodied intelligent technology

Realistic 3D embodied intelligence is an emerging research direction formed by the deep intersection of artificial intelligence, remote sensing science, and 3D reconstruction. Its core lies in transforming multi-source remote sensing observations into perceptible, inferential, and interactive 3D digital twin environments, providing intelligent agents with a spatial cognitive foundation that is highly consistent with the real world. Unlike traditional perception paradigms that rely solely on two-dimensional images or discrete point clouds, real-life 3D constructs continuous and structured environmental representations in large-scale Earth space by uniformly expressing geometric structures, semantic information, and physical attributes, enabling agents to conduct autonomous exploration and reasoning in a closed-loop mechanism of "perception cognition decision action"(朱庆 等,2022). This feature extends embodied intelligence from closed laboratory simulation environments to real and complex Earth systems, providing a unified reference frame for cross regional, cross modal, and cross task applications(Fig. 10).
figure

Fig 10 Real 3D embodied intelligence in remote sensing scenarios

With the development of remote sensing large-scale models, real-time 3D embodied intelligence is gradually shifting from "geometric reconstruction driven" to "model cognitive driven". From the perspective of infrastructure, representative remote sensing models exhibit significant differences in modeling objects and expression methods, as demonstrated by PointNet and its subsequent work(Qi等,2017)The point cloud network represented by emphasizes local global feature aggregation for unstructured 3D sampling, which is suitable for tasks such as scene classification and object recognition; NeRF and its extended models(Mildenhall等,2022)Continuous modeling of spatial radiation characteristics through implicit neural fields has shown outstanding performance in high fidelity 3D reconstruction and viewpoint synthesis, but its computational cost and dynamic modeling capabilities are still limited. In recent years, remote sensing models that integrate Transformer and multimodal encoders have emerged, which unify the processing of optics SAR、 Point cloud and timely observation data significantly enhance cross modal feature alignment and spatial semantic understanding capabilities, providing a new technological path for cognitive modeling of real-world 3D scenes. In terms of embodied intelligence theory, the early representative work "Intelligence Without Representation"(Brooks,1991)Proposed the robot concept of perception action direct connection, 《Understanding Intelligence》(Pfeifer和Scheier,2001)Systematically developed the theoretical framework of embodied intelligence. In recent years, Habitat(Savva等,2019)The platform provides a unified framework for training and evaluating embodied intelligent agents in large-scale 3D simulation environments, gradually moving embodied intelligence research from conceptual exploration to engineering practice.
Based on the above technological accumulation, the intersection of real-life 3D and embodied intelligence is gradually becoming the forefront direction of the combination of artificial intelligence and remote sensing science. In disaster response, intelligent agents can autonomously search the affected area and provide auxiliary decision-making; In urban management, embodied intelligence can actively inspect and achieve monitoring and optimization of complex systems such as transportation, energy, and environment. As a result, real-life 3D not only builds a high-precision spatial data base, but also becomes the operational "stage" for intelligent agents to carry out complex tasks; And embodied intelligence transforms remote sensing and 3D reconstruction results into high-level applications oriented towards cognition and action. However, this field still faces challenges such as cross source data fusion, dynamic scene modeling, simulation to reality migration, and model reliability. Future research can rely on dynamic neural fields, multimodal large models, and digital twin platforms to promote the widespread application of real-life 3D embodied intelligence in larger scale, multitasking scenarios(Wang等,2025a).

4. Cross application of artificial intelligence and remote sensing science

With the deep integration of artificial intelligence and remote sensing science, remote sensing applications are gradually shifting from geometric feature mapping and feature extraction of single scenes to systematic cognition and dynamic understanding of complex real-world processes. Key areas such as geology, cities, wetlands, and coastal zones have a global and strategic position in resource security, ecological environment, and sustainable development. Their observation objects generally exhibit highly dynamic changes, multi-scale coupling, and complex formation mechanisms, making them key scenarios for verifying and enhancing remote sensing intelligent perception capabilities. Based on deep learning semantic understanding, multimodal fusion, and long-term inference, remote sensing interpretation breaks through the limitations of traditional experience and single source data, achieving intelligent characterization of earth processes such as mineral resource distribution, urban expansion processes, wetland carbon sink dynamics, and coastal zone morphology evolution; At the same time, by introducing reasoning mechanisms that incorporate physical constraints and knowledge guidance, artificial intelligence is no longer limited to pixel level classification or geometric recognition, but further participates in mechanism revelation, risk warning, and decision support, gradually constructing an interpretable, transferable, and generalizable intelligent cognitive system. Based on the above progress, this chapter will systematically review the research status, key technological paths, and future development directions of artificial intelligence enabled remote sensing technology, focusing on six typical application scenarios including geological remote sensing, urban remote sensing, wetland remote sensing, ocean/coastal remote sensing, disaster remote sensing, and ecological remote sensing(Fig. 11).
figure

Fig 11 Applications of artificial intelligence in remote sensing science

4.1 Geological Remote Sensing Applications

In the field of geological remote sensing, artificial intelligence technology has gradually been deeply integrated into the core business and workflow of geological resource exploration, environmental monitoring, and earth evolution research. Firstly, multi-source remote sensing data such as visible light and shadow images, hyperspectral images, synthetic aperture radar (SAR), and polarimetric SAR images have been widely used in lithology identification, mineral prediction, soil attribute inversion, and geological structure research, achieving significant results(Janga等,2023Ma等,2019Xiao等,2025). Among them, visible light and shadow images have long been used for geological mapping and lithological zoning due to their advantages of wide coverage and low acquisition cost(Adiri等,2020); Hyperspectral remote sensing, with its fine band resolution, is used for precise mineral detection(Cao等,2020Lei等,2024Peng等,2024Zhao等,2020)Unique advantages demonstrated in trace element inversion and mixed pixel unmixing(李明 等,2021秦凯 等,2025王建华 等,2021赵宁博 等,2021); SAR and polarimetric SAR images have penetration capability and rich scattering information, which can achieve lithology identification and geological structure extraction in snow covered or vegetation covered areas.
Secondly, through artificial intelligence driven processing and analysis of high-resolution and multi-source geological remote sensing big data, applications such as fine identification of rock structural planes, inversion of Martian mineral abundance, and high-precision prediction of soil elements can be achieved. For example, using SAR combined with Consistent High Resolution Generative Adversarial Network (C-HR-GAN) can accurately identify rock structural planes in high-altitude mining areas; Based on the fully automated spectral unmixing method (FAHU), precise identification and abundance inversion of Martian water bearing minerals can be achieved, providing technical support for planetary geological evolution analysis(Ke等,2024); In terms of soil element inversion, integrating aerial hyperspectral data, physical and chemical parameters, and terrain factors, high-precision prediction of selenium and total phosphorus content in black soil can be achieved through random forest or deep neural network models(赵宁博 等,2021). In addition, the combination of polarization decomposition parameters and deep learning lithology classification methods has also achieved significant improvements in the accuracy of lithology identification in vegetation covered areas. In terms of dynamic monitoring of mines, the automation extraction of tailings pond and mining site changes is achieved through PSPED and CTMNet, providing a reliable means for mine safety supervision(Xing等,2024Zhang等,2023a).
Finally, although the combination of artificial intelligence and geological remote sensing has achieved fruitful results, it still faces challenges such as strong data heterogeneity, insufficient annotated samples, and insufficient integration of algorithms and geological mechanisms. In the future, further enhancing the processing capability of multi-source remote sensing data, promoting the transformation of artificial intelligence from data-driven to knowledge guided, and strengthening the research on basic large-scale models, it is expected to achieve systematic and large-scale development of geological remote sensing in mineral prediction, mine safety, ecological restoration, and planetary exploration applications.

4.2 Ecological Remote Sensing Applications

In the field of ecological remote sensing, land cover monitoring is the core and fundamental task, and related research has covered data acquisition, method innovation, and application practice. Through strategies such as multi-source temporal remote sensing data fusion and active learning, the limitations caused by the scarcity of annotated data can be effectively alleviated(Wu和Prasad,2018Xu等,2024b); Based on deep convolutional networks, attention mechanisms, and other methods, it is possible to collaboratively mine spectral, spatial, and temporal features, achieving accurate differentiation and dynamic perception of complex terrain types(Ma等,2026Tian等,2025a); In addition, using knowledge graph technology to correlate remote sensing information with socio-economic driving factors can support in-depth analysis and scientific decision-making of the causal mechanism of land cover change(Kafy等,2022Wang等,2025f).
The introduction of artificial intelligence has gradually expanded the research scope of ecological remote sensing from static mapping and monitoring to quantification of ecosystem functions, analysis of ecological process mechanisms, and assessment of the impact of interference events. For example, deep learning models guided by physical mechanisms exhibit higher accuracy in remote sensing estimation of total primary productivity of vegetation compared to traditional statistical or mechanistic models(Ma等,2025); Integrating multi-source remote sensing data and process mechanism models can accurately analyze the interaction between plant water use and atmospheric water demand dynamics, as well as their impact on regional climate under the background of global warming(Wang等,2025c); Based on the combination of artificial intelligence and remote sensing, monitoring of disturbances such as forest fires and grazing can achieve precise quantification of ecosystem change trajectories and their uncertainties(Chang等,2024Ma等,2020aShang等,2023).
Overall, artificial intelligence endows ecological remote sensing with disruptive analytical capabilities, achieving a leap from macroscopic observation to intelligent understanding, promoting the development of ecological remote sensing monitoring towards higher accuracy, automation, and intelligence, and providing new technological support for global ecosystem management and sustainable development.

4.3 Wetland Remote Sensing Applications

In the field of wetland remote sensing, wetlands, as highly productive ecosystems, play a key role in carbon cycle regulation, biodiversity maintenance, climate regulation, and hydrological conservation(Gardner和Finlayson,2018Rogers等,2019Tan等,2022). However, the distribution pattern of wetlands is complex, with significant dynamic changes, and some areas have poor accessibility, making it difficult for traditional ground surveys to achieve large-scale, long-term continuous monitoring(Wang等,2019王宗明 等,2025). With the advantages of multi-source, multi-scale, and long-term observation, remote sensing technology has become the core means of wetland monitoring(毛德华 等,2023Ozesmi和Bauer,2002). Wetland remote sensing refers to a methodological system for systematically identifying and analyzing wetland spatial patterns, dynamic evolution, and ecological processes based on remote sensing data(张柏,1996)The research focus has gradually expanded from early distribution mapping and type classification to ecological process characterization and key parameter inversion. In recent years, the deep integration of artificial intelligence has significantly improved the efficiency and analytical ability of remote sensing data processing, providing a new technological path for wetland research to transition from pattern cognition to functional simulation and process deduction.
In terms of data type applications, optical remote sensing images are the earliest and most widely used data source in wetland research(Mahdavi等,2018). Early aerial photography and ground surveys laid the foundation for wetland distribution mapping, while with the development of satellite remote sensing, Landsat(Zhang等,2024)The MODIS(Bansal等,2017)The Sentinel-2(Jia等,2019Li等,2019)And PlanetScope(Gonçalves等,2023)The continuous improvement of spatiotemporal resolution through multispectral data provides reliable support for large-scale wetland coverage mapping and long-term dynamic analysis.Mao 等(2020)By combining object-oriented and hierarchical classification methods (HOHC), the first high-precision national scale wetland distribution map in China was constructed, demonstrating the potential application of traditional remote sensing methods in large-scale wetland mapping. At the same time, the introduction of machine learning and deep learning methods has significantly improved the accuracy and efficiency of optical images in classification, feature extraction, and time series analysis(Tian等,2025bZhao等,20222023a). For example,Jia 等(2024)A hybrid model combining machine learning and active learning strategies was proposed to achieve high-precision inversion of multifunctional traits of mangrove forests, demonstrating the significant advantages of deep learning in optical remote sensing analysis.
Hyperspectral imaging, with its continuous bands and rich information, can capture subtle spectral differences that are difficult to distinguish with traditional multispectral imaging, significantly expanding the research depth of wetland vegetation species identification, community division, and physiological and biochemical parameter inversion(Yang等,2022Zhang等,2021). Faced with the high dimensionality and information redundancy of hyperspectral data, machine learning and deep learning methods have played a key role in feature extraction and data parsing(Gao等,2023). Lidar can accurately obtain vertical structure information of vegetation, reveal the three-dimensional distribution and community characteristics of wetland vegetation, and provide support for parameter inversion such as tree height, leaf area index, and biomass, thereby improving the accuracy of wetland carbon stock assessment(Yin和Wang,2019). Synthetic Aperture Radar (SAR) has the ability to penetrate clouds, mist, and some vegetation canopies, and has significant advantages in the humid coastal wetland environment with frequent rainfall; Combining deep convolutional networks with time series analysis methods can improve the performance of SAR in wetland vegetation classification and flood dynamic monitoring(Mahdianpari等,2017Peña等,2024). With the diversification of remote sensing sources, a single data type is no longer sufficient to meet the research needs of complex wetland ecosystems. The introduction of artificial intelligence methods has significantly improved the efficiency of multi-source data fusion and information extraction, providing favorable support for accurate wetland mapping, dynamic monitoring, and ecological process analysis. For example,Cai 等(2020)Based on SAR and optical images, an object-oriented stacked generalization method was used to achieve high-precision classification of high wetland types, fully demonstrating the technical advantages of artificial intelligence in extracting complex wetland information.
Although artificial intelligence and remote sensing technology have made significant progress in wetland monitoring, they still face various challenges. At the data level, high-quality and spatially balanced training samples are still lacking, and there is a lack of a unified annotation system across regions and wetland types, which limits the transferability of the model; At the algorithmic level, the generalization performance of deep learning models is limited, making it difficult to adapt to diverse wetland types and complex habitat conditions(Peña等,2024)At the same time, the processing of high-resolution and multi-source remote sensing data poses higher requirements for the calculation examples, and the model decision-making process also lacks interpretability(Gevaert,2022). In addition, non-technical factors such as institutional constraints, policy support, and ethical considerations may also affect the large-scale promotion of artificial intelligence in wetland monitoring in practical applications.

4.4 Agricultural Remote Sensing Applications

In the field of agricultural remote sensing, the deep integration of remote sensing technology and artificial intelligence is driving the development of research from traditional static information extraction to intelligent and dynamic comprehensive analysis and decision support(哈斯图亚和陈仲新,2025). Artificial intelligence models, with their excellent feature extraction and pattern recognition capabilities, can accurately interpret crop spectral, spatial, and temporal information from multi-source and multi-phase remote sensing data, providing key technical support for building high-precision, high-efficiency, and intelligent agricultural monitoring systems(Dou等,2024).
In terms of crop recognition, artificial intelligence algorithms integrate spectral, texture, terrain, and temporal features to achieve automatic recognition of cultivated land types and dynamic updates of planting structures, greatly improving the accuracy and efficiency of land use information extraction. For crop growth monitoring and yield prediction, artificial intelligence can analyze the complex relationship between remote sensing spectra and crop physicochemical parameters, achieve accurate estimation of key physiological parameters such as leaf area index, aboveground biomass, chlorophyll and nitrogen content, and provide yield prediction(Lu等,2023). The integration of artificial intelligence and mechanistic models (such as PROSAIL and WOFOST) has gradually become a development trend by introducing physical process constraints into data-driven models, which not only preserves the interpretability of mechanistic models, but also compensates for their shortcomings in parameter uncertainty and regional applicability(Kheir等,2025).
With the intensification of climate change, extreme events such as drought, floods, freezing damage, and pest infestations are becoming increasingly prominent threats to agricultural production. Combining remote sensing data, artificial intelligence models have demonstrated significant advantages in disaster monitoring and response. For example, regarding floods and waterlogging(Paul等,2018)Drought (Amani and Shafizadeh Moghadam, 2023), crop lodging(Tian等,2021)And pests and diseases(黄文江 等,2019)Artificial intelligence can quickly identify the scope and severity of disaster situations, and dynamically monitor and predict trends based on time-series images. With the improvement of spatial and temporal resolution of remote sensing data, as well as the maturity of intelligent interpretation algorithms, the dynamic monitoring and rapid assessment of disasters based on time-series images have shown significant spatiotemporal advantages and economic value.
In addition, in terms of agricultural resource management and intelligent decision support, artificial intelligence technology integrates meteorological, soil, moisture, and crop remote sensing data to achieve refined agricultural information management from the field to the regional scale(Chlingaryan等,2018). By combining time series analysis and prediction models, crop growth status and potential risks can be dynamically tracked, providing scientific basis for precision fertilization, irrigation management, and yield prediction, thereby supporting intelligent management and sustainable development of agricultural production(Li等,2025d).

4.5 Ocean/Coastal Remote Sensing Applications

Ocean and coastal remote sensing is an important component of global environmental monitoring. The ocean covers about two-thirds of the Earth's surface area and not only provides necessary space and resources for biological survival, but also plays a key role in global climate regulation and sustainable human development. As a transitional zone with the most intense interaction between land and sea, the coastal zone has a unique geographical form and dynamic process, which plays an important role in maintaining beach stability, protecting biological habitats, promoting material cycling, and achieving carbon sequestration and climate regulation. However, in recent years, under multiple pressures such as climate change, high-intensity human activities, and biological invasions, the ocean and coastal zones are facing severe challenges such as resource scarcity, ecological degradation, environmental pollution, and ocean warming. As a maritime power, China has approximately 3 million square kilometers of jurisdictional waters, 18000 kilometers of mainland coastline, and 14000 kilometers of island coastline. In this context, enhancing the dynamic monitoring, information acquisition, and management decision-making capabilities of the ocean/coastal zone is of great significance for promoting the strategy of building a maritime power and deep sea development. The integration of artificial intelligence and remote sensing technology provides strong technical support for scientific understanding and precise management of oceans and coastal zones.
The ocean and coastal zone are typical dynamic systems that interact with water bodies, land, atmosphere, and human activities, characterized by vast space, active energy exchange, diverse environmental factors, and significant ecological vulnerability. Related research heavily relies on the support of multi platform and multimodal remote sensing data. Artificial intelligence technology, with its powerful feature learning and information mining capabilities, has become a key means of processing multi-source heterogeneous data in the ocean and coastal zone, significantly improving the quality of data fusion, information extraction efficiency, parameter inversion accuracy, and 3D environment reconstruction capabilities.
With the development of Earth big data and multi satellite network observation technology, the data sources for remote sensing monitoring of oceans and coastal zones are becoming increasingly diversified. In response to the highly dynamic and spatiotemporal heterogeneity of the region, it is necessary to develop short-term and high-frequency continuous observation capabilities, promote multi-source multimodal remote sensing fusion, cross sensor heterogeneous data fusion, and time series remote sensing reconstruction technologies, in order to comprehensively improve observation quality. By integrating multi-source data such as ground measurements, temporal remote sensing, LiDAR, and satellite altimeters, and incorporating artificial intelligence and image processing methods, it is possible to effectively overcome technical bottlenecks such as high-precision depth measurement, high-density sampling, small footprints, and high revisit frequency, and achieve precise information extraction, target recognition, and parameter quantitative inversion. Taking ICESat-2 single photon lidar satellite as an example, it can penetrate clear water bodies and directly obtain high-precision underwater terrain and depth information in shallow water areas with photon level detection sensitivity and high-density orbital sampling capability(Ma等,2020bXu等,2024a). asFig. 12As shown, in turbid waters, combined with optical image time series, ICESat-2(Xin等,2025)It can obtain the elevation of coastal tidal flats under low tide conditions, achieve high-precision terrain inversion under complex water quality conditions, and provide a reliable technical path for wide area dynamic coastal zone monitoring(Ma等,2023Xin等,2025Xu等,2022b).
figure
figure

Fig 12 Multi-source satellite remote sensing combined with knowledge-guided deep learning for super-resolution reconstruction of three-dimensional temperature anomalies (Ⅰ) and salinity anomalies (Ⅱ) at 2000 m in the global upper ocean (Wang et al., 2024a

To address the challenges of global climate change and sustainable ocean development, it is necessary to integrate multi-source remote sensing, buoy observations, ocean models, and reanalysis data, and develop knowledge driven artificial intelligence deep-sea remote sensing technology. This technology can extend satellite observation capabilities from the surface of the sea to the middle and deep layers, achieving high-resolution remote sensing inversion and reconstruction of the ocean's three-dimensional environment, significantly improving the accuracy and resolution of ocean three-dimensional detection. On this basis, a high-quality ocean 3D observation dataset covering temperature, salinity, density, flow field, and other aspects that need to be improved can be systematically constructed, providing key data support for the study of ocean environment evolution and climate change research(Su等,2021a20242025Wang等,2024a)(Fig. 12). Through innovative artificial intelligence deep sea remote sensing methods, it is possible to reconstruct global ocean environmental change data with high precision, high resolution, long time series, and deep level features, providing a new perspective for understanding the dynamics of the ocean's sub surface and mid deep layers, effectively serving the sustainable development of the ocean and global climate assessment.

4.6 Remote sensing applications for disasters

In the field of disaster remote sensing, the rapid development of ground remote sensing observation technology provides a diversified data foundation for dynamic monitoring of disasters and extraction of disaster information(Wang等,2022a). Artificial intelligence technologies represented by deep learning have made significant progress in fields such as natural language processing and computer vision. However, due to the high complexity of disaster scenarios, their application in disaster remote sensing still faces challenges such as sample scarcity and limited model generalization ability(Wang等,2025d). Therefore, relevant research has explored from multiple perspectives based on the different stages of disaster development.
In the stage of disaster risk assessment, remote sensing technology can provide scientific basis for pre disaster warning. Evaluating disaster risk levels and regional vulnerability based on satellite remote sensing data can help support decision-making for disaster prevention and reduction(Liu等,2017). Further integrating remote sensing variables, social vulnerability indicators, and machine learning models not only improves the timeliness and accuracy of disaster warning, but also provides reliable basis for policy formulation and emergency management, thereby helping to reduce potential losses and enhance regional resilience(Samadzadegan等,2025Tan等,2020Zhang等,2023b).
In the emergency response phase of disasters, remote sensing technology provides key data support for the assessment of damage to disaster bearing bodies and emergency search and rescue(Wang等,2024c). By integrating artificial intelligence technology with multi-source satellite data before and after disasters, including satellite imagery (Gupta, etc.) and drone data(Rahnemoonfar等,2021)The system can automatically learn the visual features of the damaged disaster bearing body using thermal infrared images and limited annotated samples, and quickly generate coverage roads(Rahnemoonfar等,2023)Buildings(Chen等,2024Zheng等,2021)Waiting for artificial facilities and landslides(Wang等,2022c)Mapping the damage to natural features in disaster areas to provide scientific reference for disaster assessment and rescue decision-making.
In the post disaster recovery and reconstruction stage, the deep integration of artificial intelligence and remote sensing technology significantly enhances the dynamic assessment and monitoring capabilities of socio-economic and ecological environment restoration. Based on multi temporal remote sensing data (such as nighttime light data)(Xiao等,2024)And social media information(Eyre等,2020Markhvida等,2020)It can achieve dynamic tracking of the socio-economic vitality and ecological restoration process in disaster areas(Liu等,2024)To provide theoretical basis for policy evaluation and resource optimization allocation in reconstruction. In addition, the visual language multimodal large model(Wang等,2025d)With its semantic and spatial correlation analysis capabilities, it can provide textual decision support for long-term recovery planning in disaster areas by integrating high-resolution remote sensing images.
In the future, disaster remote sensing will present a development trend of deep integration of multiple platforms and intelligence. On the one hand, by integrating the wide coverage of satellites with the flexible and precise observation advantages of drones, a multi-level disaster information perception system can be constructed. On the other hand, multimodal remote sensing big data from space, sky, and earth will be coupled with cutting-edge artificial intelligence technology systems such as embodied intelligence and intelligent agents, promoting the intelligent upgrade of disaster management from disaster prevention, reduction to relief, and achieving overall improvement in disaster response efficiency and precision.

4.7 Urban Remote Sensing Applications

In the field of urban remote sensing, cities, as highly concentrated carriers of human activities, are key regions for understanding global change and promoting sustainable development. Urban remote sensing mainly observes surface elements, environmental processes, and human activities, relying on the advantages of large-scale coverage, multi-scale observation, and high-frequency updates, to achieve effective monitoring of urbanization process and environmental evolution(Herfort等,2023). In recent years, artificial intelligence technologies represented by deep learning have achieved deep fusion with multi-source remote sensing data with big data features (such as optical images, SAR/InSAR, hyperspectral, LiDAR, etc.), significantly improving the automation level and fine analysis level of remote sensing data(Fig. 13).
figure

Fig 13 The structure of data, methods and applications for urban remote sensing

In terms of urban feature extraction and classification, a study based on optical imaging and machine learning methods evaluated the integrity of OSM building data in global urban agglomerations, revealed significant spatial imbalances, and proposed an evaluation framework to address data bias, providing reference for urban feature extraction and sustainable development analysis(Herfort等,2023). In terms of surface deformation and 3D modeling, the fusion of optical images and LiDAR data, combined with deep learning and multi-source information fusion methods, has successfully constructed a high-precision building database at the global scale, achieving precise mapping of building 3D heights and providing a data foundation for global building form analysis(Zhu等,2025). In terms of urban environmental monitoring, artificial intelligence methods have been applied to systematically evaluate the potential of building energy conservation and carbon reduction. Research has shown that they can significantly reduce energy consumption and carbon emissions in multiple scenarios, providing quantitative basis for urban environmental management and carbon neutrality pathways(Ding等,2024). In addition, in terms of urban dynamic change detection, semantic level change recognition of urban land use/cover has been achieved through high-resolution optical imaging and deep learning frameworks, significantly improving the accuracy and generalization ability of change detection(Zhu等,2022c)。
Despite significant progress, urban remote sensing still faces many challenges. The inconsistency in spatial resolution and scale of multi-source data increases the difficulty of data fusion; In complex terrain backgrounds, the model's generalization ability and interpretability are still insufficient; Large scale applications pose significant pressure on computing and storage resources; At the same time, practical factors such as privacy protection, data sharing, and governance norms also constrain the widespread implementation of technology(Herfort等,2023). Overall, with the continuous enhancement of collaborative processing capabilities for multi-source data and breakthroughs in intelligent algorithms, artificial intelligence driven urban remote sensing will play a more systematic and scalable supporting role in environmental monitoring, socio-economic analysis, and sustainable development, providing a solid data and methodological foundation for global urban governance.

5 Challenges and Prospects

The deep integration of artificial intelligence and remote sensing science is driving the evolution of the Earth observation paradigm from "passive perception" to "active cognition". The core goal of remote sensing development has always been to achieve more accurate perception and understanding of the Earth. Although artificial intelligence technology has significantly improved the automation and accuracy level of remote sensing interpretation, existing methods still face a series of structural challenges in the real Earth system where multi-source heterogeneity, dynamic evolution, and cross scale coupling coexist. Firstly, the significant differences in imaging mechanism, spatial resolution, and temporal sampling frequency of multi-source remote sensing data significantly increase the difficulty of cross modal feature alignment and consistency modeling, limiting the robustness and generalization ability of the model in complex scenes. Secondly, existing remote sensing interpretation methods rely heavily on static sample assumptions, making it difficult to depict the continuous evolution characteristics of land cover morphology and surface processes. This limits their performance in tasks such as disaster evolution, ecological change, and urban dynamic monitoring. In addition, the dependence of deep learning models on annotated samples and regional distribution still poses significant performance degradation issues in cross regional and cross disaster applications. Finally, the lack of interpretability and physical constraints in the decision-making process of the model limits its credible application in Earth science research and high-risk decision-making scenarios. These issues have been systematically reflected in the previous chapters' discussions on multimodal modeling, remote sensing large-scale models, and embodied intelligence, and constitute key bottlenecks that need to be overcome in future research.
In response to the above challenges, the deep integration of artificial intelligence and remote sensing science is forming a technological evolution path from "perception enhancement" to "cognitive modeling". Firstly, at the model level, it is necessary to promote the collaborative modeling of general remote sensing models and physical mechanism models. By explicitly embedding radiation transmission, scattering mechanisms, and spatiotemporal constraints into deep networks, a bidirectional closed loop of data-driven and mechanism constraints can be achieved, thereby improving the stability and interpretability of the model in complex environments. Secondly, in terms of methodology, the interpretation paradigm should shift from a single task oriented approach to a knowledge driven modeling approach centered around unified representation, and build a large-scale remote sensing model with cross task and cross regional migration capabilities to alleviate sample dependence and regional limitations. Again, in the temporal dimension, the introduction of spatiotemporal models and sequence modeling mechanisms extends remote sensing observations from discrete interpretation to continuous cognitive processes, enabling the model to explicitly depict the evolutionary patterns of surface processes. Furthermore, by combining embodied intelligence with digital twin technology, intelligent agents can perform active perception, hypothesis verification, and strategy deduction in a highly consistent Earth space environment, providing more forward-looking decision support capabilities for disaster response and resource management.
At a higher level, artificial intelligence is driving the transformation of remote sensing science from a passive observation tool to an intelligent engine for the Earth system. The traditional interpretation mode centered on pixels or objects is gradually evolving into a cognitive modeling paradigm centered on geographic entities, spatial relationships, and dynamic processes, transforming land features from static targets to functional nodes that interact with each other in the Earth system. With the continuous development of multimodal large models, intelligent agent inference mechanisms, and interdisciplinary knowledge fusion, remote sensing systems are expected to achieve a leap in capabilities from information extraction and pattern recognition to causal inference and scenario prediction. This transformation will make remote sensing not only serve surface feature recognition, but also become an important intelligent infrastructure for understanding the operating mechanism of the Earth system, supporting scientific discoveries and public decision-making.
Overall, the cross fusion of artificial intelligence and remote sensing science not only represents an upgrade in the technological system, but also marks a fundamental shift in the cognitive paradigm of remote sensing. In the future, remote sensing science will no longer be limited to "observing the Earth", but will gradually move towards a new stage of "Earth System Intelligence" supported by data-driven, mechanism constrained, and intelligent decision-making through intelligent models and embodied systems to achieve structured understanding, dynamic deduction, and scientific prediction of the Earth system.

6 Conclusion

The deep integration of artificial intelligence and remote sensing science is driving fundamental changes in human cognitive patterns and application paradigms of the Earth system. From resource and environmental monitoring to disaster assessment, from urban governance to climate change system research, artificial intelligence not only provides a powerful technical engine for intelligent processing and accurate interpretation of remote sensing data, but also constructs a new path for cross scale, cross modal, and interdisciplinary comprehensive research. With the help of deep representation and knowledge transfer mechanisms, artificial intelligence promotes the transition of remote sensing from "data acquisition" to "cognitive understanding", gradually achieving intelligent analysis and process prediction of complex Earth systems. However, there are still challenges in this field such as imbalanced data, insufficient model generalization ability, and limited algorithm interpretability, which urgently require breakthroughs in theoretical models, computational frameworks, and knowledge expression. Looking ahead to the future, with the continuous development of multimodal big models, self supervised learning, causal reasoning, and physical mechanism fusion, artificial intelligence will further empower remote sensing science, achieve higher levels of automation, generalization, and intelligence, and build a solid technological foundation and cognitive framework for Earth system science, sustainable development, and global governance. It can be foreseen that the deep integration of artificial intelligence and remote sensing will not only reshape the paradigm of earth science research, but also generate broad and profound strategic value in policy formulation, social governance, and industrial transformation.

References

  1. 1.
    Adiri Z, Lhissou R, El Harti A, Jellouli A and Chakouri M. 2020. Recent advances in the use of public domain satellite imagery for mineral exploration: a review of Landsat-8 and Sentinel-2 applications. Ore Geology Reviews, 117: 103332
  2. 2.
    Agarwal S, Mustavee S, Contreras-Castillo J and Guerrero-Ibañez J. 2022. Chapter 20 - Sensing and monitoring of smart transportation systems//Alavi A H, Feng M Q, Jiao P C and Sharif-Khodaei Z, eds. The Rise of Smart Cities. Oxford: Butterworth-Heinemann: 495-522
  3. 3.
    Ahmad R. 2024. Smart remote sensing network for disaster management: an overview. Telecommunication Systems, 87(1): 213-237
  4. 4.
    Ahmadian N, Ullmann T, Verrelst J, Borg E, Zölitz R and Conrad C. 2019. Biomass assessment of agricultural crops using multi-temporal dual-polarimetric TerraSAR-X data. PFG – Journal of Photogrammetry, Remote Sensing and Geoinformation Science, 87(4): 159-175
  5. 5.
    Ali I, Greifeneder F, Stamenkovic J, Neumann M and Notarnicola C. 2015. Review of machine learning approaches for biomass and soil moisture retrievals from remote sensing data. Remote Sensing, 7(12): 16398-16421
  6. 6.
    Alkhatib M Q, Zitouni M S, Al-Saad M, Aburaed N and Al-Ahmad H. 2025. PolSAR image classification using shallow to deep feature fusion network with complex valued attention. Scientific Reports, 15(1): 24315
  7. 7.
    Amani S and Shafizadeh-Moghadam H. 2023. A review of machine learning models and influential factors for estimating evapotranspiration using remote sensing and ground-based data. Agricultural Water Management, 284: 108324
  8. 8.
    Amini S, Saber M, Rabiei-Dastjerdi H and Homayouni S. 2022. Urban land use and land cover change analysis using random forest classification of landsat time series. Remote Sensing, 14(11): 2654
  9. 9.
    An X, Sun J X, Gui Z H and He W. 2025. CHOICE: benchmarking the remote sensing capabilities of large vision-language models//The 39th Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. San Diego: [s.n.]
  10. 10.
    Bahrami H, Homayouni S, Safari A, Mirzaei S, Mahdianpari M and Reisi-Gahrouei O. 2021. Deep learning-based estimation of crop biophysical parameters using multi-source and multi-temporal remote sensing observations. Agronomy, 11(7): 1363
  11. 11.
    Bansal S, Katyal D and Garg J K. 2017. A novel strategy for wetland area extraction using multispectral MODIS data. Remote Sensing of Environment, 200: 183-205
  12. 12.
    Bochkovskiy A, Wang C Y and Liao H Y M. 2020. YOLOv4: optimal speed and accuracy of object detection. arXiv preprint arXiv: 2004.10934
  13. 13.
    Bondi E, Jain R, Aggrawal P, Anand S, Hannaford R, Kapoor A, Piavis J, Shah S, Joppa L, Dilkina B and Tambe M. 2020. BIRDSAI: a dataset for detection and tracking in aerial thermal infrared videos//Proceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision. Snowmass: IEEE: 1736-1745
  14. 14.
    Bozcan I and Kayacan E. 2020. AU-AIR: a multi-modal unmanned aerial vehicle dataset for low altitude traffic surveillance//2020 IEEE International Conference on Robotics and Automation (ICRA). Paris: IEEE: 8504-8510
  15. 15.
    Brooks R A. 1991. Intelligence without representation. Artificial Intelligence, 47(1/3): 139-159
  16. 16.
    Cai Y T, Li X Y, Zhang M and Lin H. 2020. Mapping wetland using the object-based stacked generalization method based on multi-temporal optical and SAR data. International Journal of Applied Earth Observation and Geoinformation, 92: 102164
  17. 17.
    Cao B, Yao H Y, Zhu P F and Hu Q H. 2025. Visible and clear: finding tiny objects in difference map//Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T and Varol G, eds. Computer Vision – ECCV 2024. Cham: Springer: 1-18
  18. 18.
    Cao Y, Bao N S, Liu S J, Zhao W and Li S M. 2020. Reducing moisture effects on soil organic carbon content prediction in visible and near-infrared spectra with an external parameter othogonalization algorithm. Canadian Journal of Soil Science, 100(3): 253-262
  19. 19.
    Celikkan E, Kunzmann T, Yeskaliyev Y, Itzerott S, Klein N and Herold M. 2025. WeedsGalore: a multispectral and multitemporal UAV-based dataset for crop and weed segmentation in agricultural maize fields//2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Tucson: IEEE: 4767-4777
  20. 20.
    Chang C C, Wang J, Zhao Y B, Cai T Y, Yang J L, Zhang G L, Wu X C, Otgonbayar M, Xiao X M, Xin X P and Zhang Y J. 2024. A 10-m annual grazing intensity dataset in 2015-2021 for the largest temperate meadow steppe in China. Scientific Data, 11(1): 181
  21. 21.
    Chen H R X, Song J, Han C X, Xia J S and Yokoya N. 2024. ChangeMamba: remote sensing change detection with spatiotemporal state space model. IEEE Transactions on Geoscience and Remote Sensing, 62: 4409720
  22. 22.
    Chlingaryan A, Sukkarieh S and Whelan B. 2018. Machine learning approaches for crop yield prediction and nitrogen status estimation in precision agriculture: a review. Computers and Electronics in Agriculture, 151: 61-69
  23. 23.
    Davis J W and Sharma V. 2007. Background-subtraction using contour-based fusion of thermal and visible imagery. Computer Vision and Image Understanding, 106(2/3): 162-182
  24. 24.
    Diakogiannis F I, Waldner F, Caccetta P and Wu C. 2020. ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data. ISPRS Journal of Photogrammetry and Remote Sensing, 162: 94-114
  25. 25.
    Ding C, Ke J, Levine M, Granderson J and Zhou N. 2024. Potential of artificial intelligence in reducing energy and carbon emissions of commercial buildings at scale. Nature Communications, 15(1): 5916
  26. 26.
    Ding L, Zuo X B, Hong D F, Guo H T, Lu J, Gong Z H and Bruzzone L. 2025. S2C: learning noise-resistant differences for unsupervised change detection in multimodal remote sensing images. arXiv preprint arXiv: 2502.12604
  27. 27.
    Dong P L and Chen Q. 2017. LiDAR Remote Sensing and Applications. Boca Raton: CRC Press
  28. 28.
    Dou P, Huang C L, Han W X, Hou J L, Zhang Y and Gu J. 2024. Remote sensing image classification using an ensemble framework without multiple classifiers. ISPRS Journal of Photogrammetry and Remote Sensing, 208: 190-209
  29. 29.
    Elman J L. 1990. Finding structure in time. Cognitive Science, 14(2): 179-211
  30. 30.
    Eyre R, De Luca F and Simini F. 2020. Social media usage reveals recovery of small businesses after natural hazard events. Nature Communications, 11(1): 1629
  31. 31.
    Fu H, Sun G Y, Zhang L, Zhang A Z, Ren J C, Jia X P and Li F. 2023. Three-dimensional singular spectrum analysis for precise land cover classification from UAV-borne hyperspectral benchmark datasets. ISPRS Journal of Photogrammetry and Remote Sensing, 203: 115-134
  32. 32.
    Gao Y, Xia S B, Wang P, Xi X H, Nie S and Wang C. 2025. LiDAR remote sensing meets weak supervision: concepts, methods, and perspectives. arXiv preprint arXiv: 2503.18384
  33. 33.
    Gao Y H, Zhang M M, Wang J J and Li W. 2023. Cross-scale mixing attention for multisource remote sensing data fusion and classification. IEEE Transactions on Geoscience and Remote Sensing, 61: 5507815
  34. 34.
    Gardner R C and Finlayson C. 2018. Global Wetland Outlook: State of the World’s Wetlands and Their Services to People. Gland: Secretariat of the Ramsar Convention
  35. 35.
    Gevaert C M. 2022. Explainable AI for Earth observation: a review including societal and regulatory perspectives. International Journal of Applied Earth Observation and Geoinformation, 112: 102869
  36. 36.
    Gonçalves D N, Marcato J, Carrilho A C, Acosta P R, Ramos A P M, Gomes F D G, Osco L P, da Rosa Oliveira M, Martins J A C, Damasceno G A, de Araújo M S, Li J, Roque F, de Faria Peres L, Gonçalves W N and Libonati R. 2023. Transformers for mapping burned areas in brazilian pantanal and Amazon with PlanetScope imagery. International Journal of Applied Earth Observation and Geoinformation, 116: 103151
  37. 37.
    González A, Fang Z J, Socarras Y, Serrat J, Vázquez D, Xu J L and López A M. 2016. Pedestrian detection at day/night time with visible and FIR cameras: a comparison. Sensors, 16(6): 820
  38. 38.
    Gu A and Dao T. 2024. Mamba: linear-time sequence modeling with selective state spaces//First Conference on Language Modeling. Philadelphia: [s.n.]
  39. 39.
    Guo H D and Liang D. 2024. Big Earth Data and its role in sustainability. Science Bulletin, 69(11): 1623-1627
  40. 40.
    Guo X, Lao J W, Dang B, Zhang Y Y, Yu L, Ru L X, Zhong L H, Huang Z Y, Wu K, Hu D X, He H M, Wang J, Chen J D, Yang M, Zhang Y J and Li Y S. 2024. SkySense: a multi-modal remote sensing foundation model towards universal interpretation for Earth Observation imagery//Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE: 27662-27673
  41. 41.
    Hasituya and Chen Z X. 2025. The development and prospect of agricultural remote sensing in the digital transformation of agrifood systems. National Remote Sensing Bulletin, 29(6): 1901-1917
  42. 42.
    He W, Yao Q M, Li C, Yokoya N, Zhao Q B, Zhang H Y and Zhang L P. 2022. Non-local meets global: an iterative paradigm for hyperspectral image restoration. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(4): 2089-2107
  43. 43.
    Herfort B, Lautenbach S, Porto de Albuquerque J, Anderson J and Zipf A. 2023. A spatio-temporal analysis investigating completeness and inequalities of global urban building data in OpenStreetMap. Nature Communications, 14(1): 3985
  44. 44.
    Hong D F, Zhang B, Li H, Li Y X, Yao J, Li C Y, Werner M, Chanussot J, Zipf A and Zhu X X. 2023. Cross-city matters: a multimodal remote sensing benchmark dataset for cross-city semantic segmentation using high-resolution domain adaptation networks. Remote Sensing of Environment, 299: 113856
  45. 45.
    Hong D F, Zhang B, Li X Y, Li Y X, Li C Y, Yao J, Yokoya N, Li H, Ghamisi P, Jia X P, Plaza A, Gamba P, Benediktsson J A and Chanussot J. 2024. SpectralGPT: spectral remote sensing foundation model. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8): 5227-5244
  46. 46.
    Hosseinpour-Zarnaq M, Omid M, Sarmadian F, Ghasemi-Mobtaker H, Alimardani R and Bohlol P. 2025. Exploring the capabilities of hyperspectral remote sensing for soil texture evaluation. Ecological Informatics, 90: 103336
  47. 47.
    Hu L, He W, Zhang L P and Zhang H Y. 2023a. Cross-domain meta-learning under dual-adjustment mode for few-shot hyperspectral image classification. IEEE Transactions on Geoscience and Remote Sensing, 61: 5526416
  48. 48.
    Hu L, He W, Zhang L P and Zhang H Y. 2025. TeDFL: toward text-driven few-shot learning for forest tree species classification with airborne hyperspectral images. IEEE Transactions on Geoscience and Remote Sensing, 63: 5530115
  49. 49.
    Hu M Q, Wu C, Du B and Zhang L P. 2023b. Binary change guided hyperspectral multiclass change detection. IEEE Transactions on Image Processing, 32: 791-806
  50. 50.
    Huang W J, Shi Y, Dong Y Y, Ye H C, Wu M Q, Cui B and Liu L Y. 2019. Progress and prospects of crop diseases and pests monitoring by remote sensing. Smart Agriculture, 1(4): 1-11
  51. 51.
    Janga B, Asamani G P, Sun Z H and Cristea N. 2023. A review of practical AI for remote sensing in Earth sciences. Remote Sensing, 15(16): 4112
  52. 52.
    Jia M M, Guo X X, Zhang L, Wang M, Wang W Q, Lu C Y, Zhao C P, Zhang R, Wang M, Yan H Q, Wang Z M and Verrelst J. 2024. Mapping mangrove functional traits from Sentinel-2 imagery based on hybrid models coupled with active learning strategies. International Journal of Applied Earth Observation and Geoinformation, 130: 103905
  53. 53.
    Jia M M, Wang Z M, Wang C, Mao D H and Zhang Y Z. 2019. A new vegetation index to detect periodically submerged mangrove forest using single-tide Sentinel-2 imagery. Remote Sensing, 11(17): 2043
  54. 54.
    Jia X Y, Zhu C, Li M Z, Tang W Q and Zhou W L. 2021. LLVIP: a visible-infrared paired dataset for low-light vision//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). Montreal: IEEE: 3489-3497
  55. 55.
    Jiang C C, Ren H Z, Li F G, Hong Z H, Huo H T, Zhang J Q and Xin J Y. 2025. Object detection from aerial multi-angle thermal infrared remote sensing images: dataset and method. ISPRS Journal of Photogrammetry and Remote Sensing, 228: 438-452
  56. 56.
    Jiang C C, Ren H Z, Yang H, Huo H T, Zhu P F, Yao Z Y, Li J, Sun M and Yang S H. 2024. M2FNet: multi-modal fusion network for object detection from visible and thermal infrared images. International Journal of Applied Earth Observation and Geoinformation, 130: 103918
  57. 57.
    Jiang H W, Peng M, Zhong Y J, Xie H F, Hao Z M, Lin J M, Ma X L and Hu X Y. 2022. A survey on deep learning-based change detection from high-resolution remote sensing images. Remote Sensing, 14(7): 1552
  58. 58.
    Jiang N, Wang K R, Peng X K, Yu X H, Wang Q, Xing J L, Li G R, Guo G D, Ye Q X, Jiao J B, Zhao J and Han Z J. 2023. Anti-UAV: a large-scale benchmark for vision-based UAV tracking. IEEE Transactions on Multimedia, 25: 486-500
  59. 59.
    Kafy A A, Saha M, Faisal A A, Rahaman Z A, Rahman M T, Liu D S, Fattah M A, Al Rakib A, AlDousari A E, Rahaman S N, Hasan M Z and Ahasan M A K. 2022. Predicting the impacts of land use/land cover changes on seasonal urban thermal characteristics using machine learning algorithms. Building and Environment, 217: 109066
  60. 60.
    Ke T, Zhong Y F, Song M, Wang X Y and Zhang L P. 2024. Mineral detection based on hyperspectral remote sensing imagery on Mars: from detection methods to fine mapping. ISPRS Journal of Photogrammetry and Remote Sensing, 218: 761-780
  61. 61.
    Kheir A M S, Govind A, Nangia V, El-Maghraby M A, Elnashar A, Ahmed M, Aboelsoud H, Gamal R and Feike T. 2025. Hybridization of process-based models, remote sensing, and machine learning for enhanced spatial predictions of wheat yield and quality. Computers and Electronics in Agriculture, 234: 110317
  62. 62.
    Kherimiche A, Ouahbi I and El Makkaoui K. 2024. Hyperspectral image classification using deep learning: a recent overview//2024 Sixth International Conference on Intelligent Computing in Data Sciences (ICDS). Marrakech: IEEE: 1-8
  63. 63.
    Khose S, Pal A, Agarwal A, Deepanshi, Hoffman J and Chattopadhyay P. 2025. SKYSCENES: a synthetic dataset for aerial scene understanding//Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T and Varol G, eds. Computer Vision – ECCV 2024. Cham: Springer: 19-35
  64. 64.
    Kim Y. 2014. Convolutional neural networks for sentence classification//Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Doha: Association for Computational Linguistics: 1746-1751
  65. 65.
    Kirillov A, Mintun E, Ravi N, Mao H Z, Rolland C, Gustafson L, Xiao T T, Whitehead S, Berg A C, Lo W Y, Dollár P and Girshick R. 2023. Segment anything//Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision. Paris: IEEE: 3992-4003
  66. 66.
    Kumar P, Prasad R, Gupta D K, Mishra V N, Vishwakarma A K, Yadav V P, Bala R, Choudhary A and Avtar R. 2018. Estimation of winter wheat crop growth parameters using time series Sentinel-1A SAR data. Geocarto International, 33(9): 942-956
  67. 67.
    Lee H, Hong S, Song J, Cho H, Jin Z X, Kim B, Jin J, Im J, Noh B and Yeo H. 2025a. DRIFT open dataset: a drone-derived intelligence for traffic analysis in urban environment. arXiv preprint arXiv: 2504.11019
  68. 68.
    Lee J, Miyanishi T, Kurita S, Sakamoto K, Azuma D, Matsuo Y and Inoue N. 2025b. CityNav: a large-scale dataset for real-world aerial navigation. arXiv preprint arXiv: 2406.14240v3
  69. 69.
    Lei H M, Bao N S, Peng S H, Yang X Y and Lu Z W. 2024. Quantitative characterization of bidirectional reflectance distribution of mine soil using physical models. European Journal of Soil Science, 75(6): e70003
  70. 70.
    Li D R, Wang M, Guo H N and Jin W J. 2025a. On China’s earth observation system: mission, vision and application. Geo-Spatial Information Science, 28(2): 303-321
  71. 71.
    Li H Y, Jia M M, Zhang R, Ren Y X and Wen X. 2019. Incorporating the plant phenological trajectory into mangrove species mapping with dense time series Sentinel-2 imagery and the Google Earth Engine platform. Remote Sensing, 11(21): 2479
  72. 72.
    Li J D, Wang Y and Qu H C. 2024. Transformer-based cross scale interactive learning for camouflage object detection. Computer Systems and Applications, 33(2): 115-124
  73. 73.
    Li J L and Wu C. 2024. Using difference features effectively: a multi-task network for exploring change areas and change moments in time series remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing, 218: 487-505
  74. 74.
    Li J P, He W, Cao W N, Zhang L P and Zhang H Y. 2024a. UANet: an uncertainty-aware network for building extraction from remote sensing images. IEEE Transactions on Geoscience and Remote Sensing, 62: 5608513
  75. 75.
    Li J P, He W, Li Z H, Guo Y J and Zhang H Y. 2025b. Overcoming the uncertainty challenges in detecting building changes from remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing, 220: 1-17
  76. 76.
    Li J P, Wei Y P, Wei T G and He W. 2025d. A comprehensive deep-learning framework for fine-grained farmland mapping from high-resolution images. IEEE Transactions on Geoscience and Remote Sensing, 63: 5601215
  77. 77.
    Li J T, Liu Y Y, Wang X Y, Peng Y N, Sun C, Wang S Y, Sun Z D, Ke T, Jiang X, Lu T W, Zhao A R and Zhong Y F. 2025c. HyperFree: a channel-adaptive and tuning-free foundation model for hyperspectral remote sensing imagery//Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE: 23048-23058
  78. 78.
    Li K, Wan G, Cheng G, Meng L Q and Han J W. 2020. Object detection in optical remote sensing images: a survey and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing, 159: 296-307
  79. 79.
    Li L Y, Huang H G, Mu X H, Yan G J, Qi J B, Yan Z B, Jiang J L, Yang H and Xiao Q. 2025. Low-altitude UAV-based quantitative remote sensing of vegetation: advances, challenges, and prospects. National Remote Sensing Bulletin, 29(6): 2083-2113
  80. 80.
    Li M, Zhu L, Yang Y C, Qin K, Zhang D H and Zhao Y J. 2021. Prediction of total phosphorus content using deep neural network and thermal infrared imaging hyperspectral data in black soil. Chinese Journal of Soil Science, 52(6): 1273-1280
  81. 81.
    Li Z H, Lu F X, Zou J Q, Hu L and Zhang H Y. 2024b. Generalized few-shot meets remote sensing: discovering novel classes in land cover mapping via hybrid semantic segmentation framework//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Seattle: IEEE: 2744-2754
  82. 82.
    Li Z L, Tang B H, Wu H, Zhao W, Duan S B, Ren H Z, Zhao E Y, Tang R L, Si M L, Leng P, Liu X Y, Liu M, Ru C, Jiang Y Z, Yan G J and Gao C X. 2025. Development and prospects of thermal infrared remote sensing. National Remote Sensing Bulletin, 29(6): 1529-1550
  83. 83.
    Liu C Y, Zhao R, Chen J Q, Qi Z P, Zou Z X and Shi Z W. 2023a. A decoupling paradigm with prompt learning for remote sensing image change captioning. IEEE Transactions on Geoscience and Remote Sensing, 61: 5622018
  84. 84.
    Liu H T, Li C Y, Wu Q Y and Lee Y J. 2023b. Visual instruction tuning//37th Conference on Neural Information Processing Systems. New Orleans: [s.n.]
  85. 85.
    Liu L, Sun X J, Chen F, Zhao S J and Gao T C. 2011. Cloud classification based on structure features of infrared images. Journal of Atmospheric and Oceanic Technology, 28(3): 410-417
  86. 86.
    Liu M W, Zhu J, Zhu Q, Qi H, Yin L Z, Zhang X, Feng B, He H G, Yang W J and Chen L Y. 2017. Optimization of simulation and visualization analysis of dam-failure flood disaster for diverse computing systems. International Journal of Geographical Information Science, 31(9): 1891-1906
  87. 87.
    Liu Q, He Z Y, Li X and Zheng Y. 2020. PTB-TIR: a thermal infrared pedestrian tracking benchmark. IEEE Transactions on Multimedia, 22(3): 666-675
  88. 88.
    Liu S B, Zhang H S, Qi Y K, Wang P, Zhang Y N and Wu Q. 2023c. AerialVLN: vision-and-language navigation for UAVs//Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision. Paris: IEEE: 15338-15348
  89. 89.
    Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C Y and Berg A C. 2016. SSD: single shot MultiBox detector//Leibe B, Matas J, Sebe N and Welling M, eds. Computer Vision – ECCV 2016. Cham: Springer: 21-37
  90. 90.
    Liu X G, Dai C G, Ding L, Zhang Z C, Li Y J, Zuo X B, Li M M, Wang H Y and Miao Y Z. 2025. GSTM-SCD: graph-enhanced spatio-temporal state space model for semantic change detection in multi-temporal remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing, 230: 73-91
  91. 91.
    Liu Y H, Lin Y, Liu W Y, Zhou J and Wang J. 2024. Remote sensing perspective in exploring the spatiotemporal variation characteristics and post-disaster recovery of ecological environment quality, a case study of the 2010 Ms7.1 Yushu earthquake. Geomatics, Natural Hazards and Risk, 15(1): 2314578
  92. 92.
    Long J, Shelhamer E and Darrell T. 2015. Fully convolutional networks for semantic segmentation//Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE: 3431-3440
  93. 93.
    Lu X P, Wang X X and Yang Z N. 2023. Leaf area index estimation from the time-series SAR data using the AIEM-MWCM model. International Journal of Digital Earth, 16(2): 4385-4403
  94. 94.
    Luo H, Feng X B, Du B and Zhang Y X. 2024. A multimodal feature fusion network for building extraction with very high-resolution remote sensing image and LiDAR data. IEEE Transactions on Geoscience and Remote Sensing, 62: 5621819
  95. 95.
    Lv P Y, Gao Y, Hu H, Cheng P and Zhong Y F. 2025. QSCDNet: a hybrid quantum spectral change detection network for hyperspectral image change detection. IEEE Transactions on Geoscience and Remote Sensing, 63: 5521212
  96. 96.
    Lyu Y, Vosselman G, Xia G S, Yilmaz A and Yang M Y. 2020. UAVid: a semantic segmentation dataset for UAV imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 165: 108-119
  97. 97.
    Ma H, Wu B F, Tian F Y, Zhao H, Li Y F, Li M X, Zhang M, Wu H T, Zeng H W and Zhu L. 2026. A deep learning-based framework for extracting spatial patterns of farmland shelterbelts in the three-north region of China using Sentinel-2 data. Remote Sensing of Environment, 332: 115095
  98. 98.
    Ma L, Liu Y, Zhang X L, Ye Y X, Yin G F and Johnson B A. 2019. Deep learning in remote sensing applications: a meta-analysis and review. ISPRS Journal of Photogrammetry and Remote Sensing, 152: 166-177
  99. 99.
    Ma Q, Bales R C, Rungee J, Conklin M H, Collins B M and Goulden M L. 2020a. Wildfire controls on evapotranspiration in california’s sierra nevada. Journal of Hydrology, 590: 125364
  100. 100.
    Ma X P, Wu Q Q, Zhao X Y, Zhang X K, Pun M O and Huang B. 2024. SAM-assisted remote sensing imagery semantic segmentation with object and boundary constraints. IEEE Transactions on Geoscience and Remote Sensing, 62: 5636916
  101. 101.
    Ma Y, Wang L, Xu N, Zhang S Y, Wang X H and Li S. 2023. Estimating coastal slope of sandy beach from ICESat-2: a case study in Texas. Environmental Research Letters, 18(4): 044039
  102. 102.
    Ma Y, Xu N, Liu Z, Yang B S, Yang F L, Wang X H and Li S. 2020b. Satellite-derived bathymetry using the ICESat-2 lidar and Sentinel-2 imagery datasets. Remote Sensing of Environment, 250: 112047
  103. 103.
    Ma Y M, Guan X B, Wang Y C, Li Y Y, Lin D K and Shen H F. 2025. GPP estimation by transfer learning with combined solar-induced chlorophyll fluorescence and eddy covariance data. International Journal of Applied Earth Observation and Geoinformation, 139: 104503
  104. 104.
    Magruder L A, Farrell S L, Neuenschwander A, Duncanson L, Csatho B, Kacimi S and Fricker H A. 2024. Monitoring Earth’s climate variables with satellite laser altimetry. Nature Reviews Earth and Environment, 5(2): 120-136
  105. 105.
    Mahdavi S, Salehi B, Granger J, Amani M, Brisco B and Huang W M. 2018. Remote sensing for wetland classification: a comprehensive review. GIScience and Remote Sensing, 55(5): 623-658
  106. 106.
    Mahdianpari M, Salehi B, Mohammadimanesh F and Motagh M. 2017. Random forest wetland classification using ALOS-2 L-band, RADARSAT-2 C-band, and TerraSAR-X imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 130: 13-31
  107. 107.
    Mao D H, Wang Z M, Du B J, Li L, Tian Y L, Jia M M, Zeng Y, Song K S, Jiang M and Wang Y Q. 2020. National wetland mapping in China: a new product resulting from object-based and hierarchical classification of Landsat 8 OLI images. ISPRS Journal of Photogrammetry and Remote Sensing, 164: 11-25
  108. 108.
    Mao D H, Wang Z M, Jia M M, Luo L, Niu Z G, Jiang W G and Sun W W. 2023. Review of global studies on the remote sensing of wetlands from 1975 to 2020. National Remote Sensing Bulletin, 27(6): 1270-1280
  109. 109.
    Markhvida M, Walsh B, Hallegatte S and Baker J. 2020. Quantification of disaster impacts through household well-being losses. Nature Sustainability, 3(7): 538-547
  110. 110.
    Maulik U and Chakraborty D. 2017. Remote Sensing Image Classification: a survey of support-vector-machine-based advanced techniques. IEEE Geoscience and Remote Sensing Magazine, 5(1): 33-52
  111. 111.
    Mildenhall B, Srinivasan P P, Tancik M, Barron J T, Ramamoorthi R and Ng R. 2022. NeRF: representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1): 99-106
  112. 112.
    Nanda A, Cho S W, Lee H and Park J H. 2024. KOLOMVERSE: korea open large-scale image dataset for object detection in the maritime universe. IEEE Transactions on Intelligent Transportation Systems, 25(12): 20832-20840
  113. 113.
    Niu R G, Sun X, Tian Y, Diao W H, Chen K Q and Fu K. 2022. Hybrid multiple attention network for semantic segmentation in aerial images. IEEE Transactions on Geoscience and Remote Sensing, 60: 5603018
  114. 114.
    Ozesmi S L and Bauer M E. 2002. Satellite remote sensing of wetlands. Wetlands Ecology and Management, 10(5): 381-402
  115. 115.
    Pande C B and Moharir K N. 2023. Application of hyperspectral remote sensing role in precision farming and sustainable agriculture under climate change: a review//Pande C B, Moharir K N, Singh S K, Pham Q B and Elbeltagi A, eds. Climate Change Impacts on Natural Resources, Ecosystems and Agricultural Systems. Cham: Springer: 503-520
  116. 116.
    Paul A, Tripathi D and Dutta D. 2018. Application and comparison of advanced supervised classifiers in extraction of water bodies from remote sensing images. Sustainable Water Resources Management, 4(4): 905-919
  117. 117.
    Peña F J, Hübinger C, Payberah A H and Jaramillo F. 2024. DEEPAQUA: semantic segmentation of wetland water surfaces with SAR imagery using deep neural networks without manually annotated data. International Journal of Applied Earth Observation and Geoinformation, 126: 103624
  118. 118.
    Peng S H, Bao N S, Wang S J, Gholizadeh A, Saberioon M and Peng Y. 2024. Mapping vertical distribution of SOC and TN in reclaimed mine soils using point and imaging spectroscopy. Ecological Indicators, 158: 111437
  119. 119.
    Petersson H, Gustafsson D and Bergstrom D. 2016. Hyperspectral image analysis using deep learning — A review//2016 Sixth International Conference on Image Processing Theory, Tools and Applications (IPTA). Oulu: IEEE: 1-6
  120. 120.
    Pfeifer R and Scheier C. 2001. Understanding Intelligence. Cambridge: MIT Press.
  121. 121.
    Portmann J, Lynen S, Chli M and Siegwart R. 2014. People detection and tracking from aerial thermal views//2014 IEEE International Conference on Robotics and Automation (ICRA). Hong Kong, China: IEEE: 1794-1800
  122. 122.
    Qi C R, Su H, Kaichun M and Guibas L J. 2017. PointNet: deep learning on point sets for 3D classification and segmentation//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE: 77-85
  123. 123.
    Qin K, Hao Y X, Zhao Y J, Cui X, Yang Y C, Zhu L and Tian Q L. 2025. A survey on hyperspectral remote sensing unmixing techniques based on autoencoders (inner cover paper·invited). Infrared and Laser Engineering, 54(5): 20250131
  124. 124.
    Rahnemoonfar M, Chowdhury T and Murphy R. 2023. RescueNet: a high resolution UAV semantic segmentation dataset for natural disaster damage assessment. Scientific Data, 10(1): 913
  125. 125.
    Rahnemoonfar M, Chowdhury T, Sarkar A, Varshney D, Yari M and Murphy R R. 2021. Floodnet: a high resolution aerial imagery dataset for post flood scene understanding. IEEE Access, 9: 89644-89654
  126. 126.
    Razakarivony S and Jurie F. 2016. Vehicle detection in aerial imagery : a small target detection benchmark. Journal of Visual Communication and Image Representation, 34: 187-203
  127. 127.
    Redmon J, Divvala S, Girshick R and Farhadi A. 2016. You only look once: unified, real-time object detection//Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 779-788
  128. 128.
    Ren S Q, He K M, Girshick R and Sun J. 2017. Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6): 1137-1149
  129. 129.
    Robinson C, Hou L, Malkin K, Soobitsky R, Czawlytko J, Dilkina B and Jojic N. 2019. Large scale high-resolution land cover mapping with multi-resolution data//Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE: 12718-12727
  130. 130.
    Rogers K, Kelleway J J, Saintilan N, Megonigal J P, Adams J B, Holmquist J R, Lu M, Schile-Beers L, Zawadzki A, Mazumder D and Woodroffe C D. 2019. Wetland carbon storage controlled by millennial-scale variation in relative sea-level rise. Nature, 567(7746): 91-95
  131. 131.
    Ronneberger O, Fischer P and Brox T. 2015. U-Net: convolutional networks for biomedical image segmentation//18th International Conference on Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. Munich: Springer: 234-241
  132. 132.
    Samadzadegan F, Toosi A and Dadrass Javan F. 2025. A critical review on multi-sensor and multi-platform remote sensing data fusion approaches: current status and prospects. International Journal of Remote Sensing, 46(3): 1327-1402
  133. 133.
    Savva M, Kadian A, Maksymets O, Zhao Y L, Wijmans E, Jain B, Straub J, Liu J, Koltun V, Malik J, Parikh D and Batra D. 2019. Habitat: a platform for embodied AI research//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE: 9338-9346
  134. 134.
    Shang R, Chen J M, Xu M Z, Lin X D, Li P, Yu G R, He N P, Xu L, Gong P, Liu L Y, Liu H and Jiao W Z. 2023. China’s current forest age structure will lead to weakened carbon sinks in the near future. The Innovation, 4(6): 100515
  135. 135.
    Shi C J, Ren B J, Wang Z W, Yan J W and Shi Z. 2022. Survey of camouflaged object detection based on deep learning. Journal of Frontiers of Computer Science and Technology, 16(12): 2734-2751
  136. 136.
    Shi J F, Ji S S, Jin H Y, Zhang Y L, Gong M G and Lin W S. 2025. Multi-feature lightweight DeeplabV3+ network for polarimetric SAR image classification with attention mechanism. Remote Sensing, 17(8): 1422
  137. 137.
    Su H, Teng J C, Zhang F Y, Wang A and Huang Z C. 2025. Can satellite observations detect global ocean heat content change with high resolution by deep learning? ISPRS Journal of Photogrammetry and Remote Sensing, 225: 52-68
  138. 138.
    Su H, Wang A, Zhang T Y, Qin T, Du X P and Yan X H. 2021a. Super-resolution of subsurface temperature field from remote sensing observations based on machine learning. International Journal of Applied Earth Observation and Geoinformation, 102: 102440
  139. 139.
    Su H, Zhang F Y, Teng J C, Wang A and Huang Z C. 2024. Reconstructing high-resolution subsurface temperature of the global ocean using deep forest with combined remote sensing and in situ observations. ISPRS Journal of Photogrammetry and Remote Sensing, 218: 389-404
  140. 140.
    Su H, Zhang T Y, Lin M J, Lu W F and Yan X H. 2021b. Predicting subsurface thermohaline structure from remote sensing data based on long short-term memory neural networks. Remote Sensing of Environment, 260: 112465
  141. 141.
    Sun G Y, Pan Z J, Zhang A Z, Jia X P, Ren J C, Fu H and Yan K. 2023a. Large kernel spectral and spatial attention networks for hyperspectral image classification. IEEE Transactions on Geoscience and Remote Sensing, 61: 5519915
  142. 142.
    Sun X, Wang P J, Lu W X, Zhu Z C, Lu X N, He Q B, Li J X, Rong X E, Yang Z J, Chang H, He Q L, Yang G, Wang R P, Lu J W and Fu K. 2023b. RingMo: a remote sensing foundation model with masked image modeling. IEEE Transactions on Geoscience and Remote Sensing, 61: 5612822
  143. 143.
    Sun Y M, Cao B, Zhu P F and Hu Q H. 2022. Drone-based RGB-infrared cross-modality vehicle detection via uncertainty-aware learning. IEEE Transactions on Circuits and Systems for Video Technology, 32(10): 6700-6713
  144. 144.
    Suo J S, Wang T Y, Zhang X Z, Chen H Y, Zhou W and Shi W S. 2023. HIT-UAV: a high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detection. Scientific Data, 10(1): 227
  145. 145.
    Tan L S, Ge Z M, Ji Y H, Lai D Y F, Temmerman S, Li S H, Li X Z and Tang J W. 2022. Land use and land cover changes in coastal and inland wetlands cause soil carbon and nitrogen loss. Global Ecology and Biogeography, 31(12): 2541-2563
  146. 146.
    Tan Q L, Wang P, Hu J, Zhou P G, Bai M Z and Hu J P. 2020. RETRACTED: the application of multi-sensor target tracking and fusion technology to the comprehensive early warning information extraction of landslide multi-point monitoring data. Measurement, 166: 108044
  147. 147.
    Tang Q, Su C, Tian Y, Zhao S B, Yang K, Hao W, Feng X B and Xie M L. 2025. YOLO-SS: optimizing YOLO for enhanced small object detection in remote sensing imagery. The Journal of Supercomputing, 81(1): 303
  148. 148.
    Tian F Y, Wu B F, Zeng H W, Zhang M, Zhu W W, Yan N N, Lu Y M and Li Y F. 2025a. GMIE: a global maximum irrigation extent and central pivot irrigation system dataset derived via irrigation performance during drought stress and deep learning methods. Earth System Science Data, 17(3): 855-880
  149. 149.
    Tian J Y, Wang L, Diao C Y, Zhang Y M, Jia M M, Zhu L, Xu M, Li X J and Gong H L. 2025b. National scale sub-meter mangrove mapping using an augmented border training sample method. ISPRS Journal of Photogrammetry and Remote Sensing, 220: 156-171
  150. 150.
    Tian M L, Ban S T, Yuan T, Ji Y B, Ma C and Li L Y. 2021. Assessing rice lodging using UAV visible and multispectral image. International Journal of Remote Sensing, 42(23): 8840-8857
  151. 151.
    Varga L A, Kiefer B, Messmer M and Zell A. 2022. SeaDronesSee: a maritime benchmark for detecting humans in open water//Proceedings of the 2022 IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa: IEEE: 3686-3696
  152. 152.
    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser Ł and Polosukhin I. 2017. Attention is all you need//Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach: Curran Associates Inc.: 6000-6010
  153. 153.
    Veličković P, Cucurull G, Casanova A, Romero A, Liò P and Bengio Y. 2018. Graph attention networks. arXiv preprint arXiv: 1710.10903
  154. 154.
    Wang A, Su H, Huang Z C and Yan X H. 2024a. Knowledge-informed deep learning model for subsurface thermohaline reconstruction from satellite observations. IEEE Transactions on Geoscience and Remote Sensing, 62: 4213416
  155. 155.
    Wang C, Yang X B, Xi X H, Nie S and Dong P L. 2024b. LiDAR remote sensing principles//Wang C, Yang X B, Xi X H, Nie S and Dong P L, eds. Introduction to Lidar Remote Sensing. Boca Raton: CRC Press
  156. 156.
    Wang D, Hu M Q, Jin Y, Miao Y C, Yang J Q, Xu Y C, Qin X L, Ma J Q, Sun L Y, Li C X, Fu C, Chen H R X, Han C X, Yokoya N, Zhang J, Xu M Q, Liu L, Zhang L F, Wu C, Du B, Tao D C and Zhang L P. 2025a. HyperSIGMA: hyperspectral intelligence comprehension foundation model. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(8): 6427-6444
  157. 157.
    Wang D, Zhang Q M, Xu Y F, Zhang J, Du B, Tao D C and Zhang L P. 2023. Advancing plain vision transformer toward remote sensing foundation model. IEEE Transactions on Geoscience and Remote Sensing, 61: 5607315
  158. 158.
    Wang H F, Cao H L, Kai Y, Bai H C, Chen X F, Yang Y, Xing L and Zhou C J. 2022a. Multi-source remote sensing intelligent characterization technique-based disaster regions detection in high-altitude mountain forest areas. IEEE Geoscience and Remote Sensing Letters, 19: 3512905
  159. 159.
    Wang H F, He W, Li Z H and Yokoya N. 2025b. Cross-scenario damaged building extraction network: methodology, application, and efficiency using single-temporal HRRS imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 228: 228-248
  160. 160.
    Wang J D, Sun K, Cheng T H, Jiang B R, Deng C R, Zhao Y, Liu D, Mu Y D, Tan M K, Wang X G, Liu W Y and Xiao B. 2021. Deep high-resolution representation learning for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10): 3349-3364
  161. 161.
    Wang J H, Zuo L, Li Z Z, Mu H Y, Zhou P, Yang J J, Zhao Y J and Qin K. 2021. A detection method of trace metal elements in black soil based on hyperspectral technology: geological implications. Journal of Geomechanics, 27(3): 418-429
  162. 162.
    Wang J J, Xuan W H, Qi H L, Liu Z H, Liu K Y, Wu Y H, Chen H R X, Song J, Xia J S, Zheng Z and Yokoya N. 2025d. DisasterM3: a remote sensing vision-language dataset for disaster damage assessment and response//The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. San Diego: [s.n.]
  163. 163.
    Wang J P, Niu H L, Zhang S P, Chen X Z, Xia X S, Zhang Y W, Lu X J, He B, Wu T W, Song C Q, Fu Z, Yao J Y and Yuan W P. 2025c. Higher warming rate in global arid regions driven by decreased ecosystem latent heat under rising vapor pressure deficit from 1981 to 2022. Agricultural and Forest Meteorology, 371: 110622
  164. 164.
    Wang L, Jia M M, Yin D M and Tian J Y. 2019. A review of remote sensing for mangrove forests: 1956-2018. Remote Sensing of Environment, 231: 111223
  165. 165.
    Wang L B, Li R, Zhang C, Fang S H, Duan C X, Meng X L and Atkinson P M. 2022b. UNetFormer: a UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 190: 196-214
  166. 166.
    Wang S, Li S, Zhang Y, Yu S, Yuan S, She R, Guo Q, Zheng J, Howe O K, Chandra L, Srijeyan S, Sivadas A, Aggarwal T, Liu H, Zhang H, Chen C, Jiang J, Xie L and Tay W P. 2025e. UAVScenes: a multi-modal dataset for UAVs//Proceedings of the IEEE/CVF International Conference on Computer Vision. 28946-28958
  167. 167.
    Wang X T, He H L, Zhang M Y, Deng J M, Ren X L, Lv Y, Liu W H, Lin Z N and Dong S Y. 2025f. Grain for Green Project dominates greening in afforested areas rather than that in grass revegetation areas of the Loess Plateau, China—using Deep Crossing LSTM Age network. Environmental Research Letters, 20(8): 084068
  168. 168.
    Wang Y Q, Dong J, Zhang L, Zhang L, Deng S H, Zhang G K, Liao M S and Gong J Y. 2022c. Refined InSAR tropospheric delay correction for wide-area landslide identification and monitoring. Remote Sensing of Environment, 275: 113013
  169. 169.
    Wang Z J, Yi J Z, Chen A B, Chen L J, Lin H and Xu K. 2025g. Accurate semantic segmentation of very high-resolution remote sensing images considering feature state sequences: from benchmark datasets to urban applications. ISPRS Journal of Photogrammetry and Remote Sensing, 220: 824-840
  170. 170.
    Wang Z J, Yi J Z, Chen A B and Han G J. 2025h. MoVis: when 3D object detection is like human monocular vision. IEEE Transactions on Image Processing, 34: 3025-3040
  171. 171.
    Wang Z J, Yi J Z, Yuan J, Hu R L, Peng X J, Chen A B and Shen X H. 2024c. Lightning-generated Whistlers recognition for accurate disaster monitoring in China and its surrounding areas based on a homologous dual-feature information enhancement framework. Remote Sensing of Environment, 304: 114021
  172. 172.
    Wang Z M, Yan Y Y, Zhao C P, Jia M M, Zhang R, Guo X X, Cheng L N, Feng Z J, Zhang Y and Chen F. 2025. Advances and perspectives in coastal wetland remote sensing research. National Remote Sensing Bulletin, 29(6): 1938-1962
  173. 173.
    Wen L Y, Du D W, Zhu P F, Hu Q H, Wang Q L, Bo L F and Lyu S. 2021. Detection, tracking, and counting meets drones in crowds: a benchmark//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE: 7808-7817
  174. 174.
    Wu H and Prasad S. 2018. Semi-supervised deep learning using pseudo labels for hyperspectral image classification. IEEE Transactions on Image Processing, 27(3): 1259-1270
  175. 175.
    Wu Z F, Cao Z, Zheng Z H, Zhang Q F, Huang X J, Liu G Y, Tan X J, Guo Y F and Li J Y. 2025. A review of urban remote sensing in China. National Remote Sensing Bulletin, 29(6): 2188-2215
  176. 176.
    Xiao A R, Xuan W H, Wang J J, Huang J X, Tao D C, Lu S J and Yokoya N. 2025. Foundation models for remote sensing and Earth observation: a survey. IEEE Geoscience and Remote Sensing Magazine, 13(4): 297-324
  177. 177.
    Xiao Z Y, Pan Y Y, Jiang L L, Wang Z and Shi K F. 2024. Remote sensing nighttime lights reveal the post-earthquake losses and reconstruction situations in Turkey–Syria earthquake areas. IEEE Geoscience and Remote Sensing Letters, 21: 3002405
  178. 178.
    Xie B, Zhang C, Wang F, Liu P, Lu F, Chen Z and Hu W. 2025. CST anti-UAV: a thermal infrared benchmark for tiny UAV tracking in complex scenes. Proceedings of the IEEE/CVF International Conference on Computer Vision. 6157-6166 [DOI:
  179. 179.
    Xie Q H, Wang J F, Lopez-Sanchez J M, Peng X, Liao C H, Shang J L, Zhu J J, Fu H Q and Ballester-Berman J D. 2021. Crop height estimation of corn from multi-year RADARSAT-2 polarimetric observables using machine learning. Remote Sensing, 13(3): 392
  180. 180.
    Xin H C, Xu N, Xu H, Yang H T, Wang Z J, Zhang Z H, Ding Y, Luan H, Ou Y F and Yang Y B. 2025. Mapping tidal flat topography by combining ICESat-2 laser altimetry and multi-source satellite imagery. International Journal of Digital Earth, 18(2): 2554313
  181. 181.
    Xing J H, Zhang J, Li J, Gao Y S, Du S H, Zhang C Y and Wang Y H. 2024. CTMNet: enhanced open-pit mine extraction and change detection with a hybrid CNN-transformer multitask network. IEEE Transactions on Geoscience and Remote Sensing, 62: 5409619
  182. 182.
    Xu C, Ding Y L, Zheng X M, Wang Y Q, Zhang R, Zhang H Y, Dai Z W and Xie Q Y. 2022a. A comprehensive comparison of machine learning and feature selection methods for maize biomass estimation using Sentinel-1 SAR, Sentinel-2 vegetation indices, and biophysical variables. Remote Sensing, 14(16): 4083
  183. 183.
    Xu N, Ma Y, Yang J, Wang X H, Wang Y J and Xu R. 2022b. Deriving tidal flat topography using ICESat-2 laser altimetry and Sentinel-2 imagery. Geophysical Research Letters, 49(2): e2021GL096813
  184. 184.
    Xu N, Wang L, Zhang H S, Tang S L, Mo F and Ma X. 2024a. Machine learning based estimation of coastal bathymetry from ICESat-2 and Sentinel-2 data. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 17: 1748-1755
  185. 185.
    Xu R T, Wang C W, Zhang J G, Xu S B, Meng W L and Zhang X P. 2023. RSSFormer: foreground saliency enhancement for remote sensing land-cover segmentation. IEEE Transactions on Image Processing, 32: 1052-1064
  186. 186.
    Xu Y J, Zhou J and Zhang Z. 2024b. A new Bayesian semi-supervised active learning framework for large-scale crop mapping using Sentinel-2 imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 209: 17-34
  187. 187.
    Yang B S and Dong Z. 2019. Progress and perspective of point cloud intelligence. Acta Geodaetica et Cartographica Sinica, 48(12): 1575-1585
  188. 188.
    Yang G, Huang K, Sun W W, Meng X C, Mao D H and Ge Y. 2022. Enhanced mangrove vegetation index based on hyperspectral images for mapping mangrove. ISPRS Journal of Photogrammetry and Remote Sensing, 189: 236-254
  189. 189.
    Yang X, Yang J R, Yan J C, Zhang Y, Zhang T F, Guo Z, Sun X and Fu K. 2019. SCRDet: towards more robust detection for small, cluttered and rotated objects//Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision. Seoul: IEEE: 8231-8240
  190. 190.
    Yao F L, Lu W X, Yang H M, Xu L Y, Liu C L, Hu L Y, Yu H F, Liu N Y, Deng C B, Tang D K, Chen C S, Yu J Q, Sun X and Fu K. 2023. RingMo-sense: remote sensing foundation model for spatiotemporal prediction via spatiotemporal evolution disentangling. IEEE Transactions on Geoscience and Remote Sensing, 61: 5620821
  191. 191.
    Yin D M and Wang L. 2019. Individual mangrove tree measurement using UAV-based LiDAR data: possibilities and challenges. Remote Sensing of Environment, 223: 34-49
  192. 192.
    Yin H T, Wang H and Zhu Z Y. 2025. Progressive dynamic queries reformation-based DETR for remote sensing object detection. IEEE Geoscience and Remote Sensing Letters, 22: 6003705
  193. 193.
    Yu H, Kong B, Hou Y T, Xu X Y, Chen T and Liu X M. 2022. A critical review on applications of hyperspectral remote sensing in crop monitoring. Experimental Agriculture, 58: e26
  194. 194.
    Zhang B. 1996. Application of remote sensing technology on research of the wetland in China. Remote Sensing Technology and Application, 11(1): 67-71
  195. 195.
    Zhang C Y, Xing J H, Li J, Du S H and Qin Q M. 2023a. A new method for the extraction of tailing ponds from very high-resolution remotely sensed images: PSVED. International Journal of Digital Earth, 16(1): 2681-2703
  196. 196.
    Zhang R, Jia M M, Wang Z M, Zhou Y M, Wen X, Tan Y and Cheng L N. 2021. A comparison of Gaofen-2 and Sentinel-2 imagery for mapping mangrove forests using object-oriented analysis and random forest. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14: 4185-4193
  197. 197.
    Zhang T J, Wang D L and Lu Y. 2023b. Machine learning-enabled regional multi-hazards risk assessment considering social vulnerability. Scientific Reports, 13(1): 13405
  198. 198.
    Zhang X, Liu L Y, Zhao T T, Wang J Q, Liu W D and Chen X D. 2024. Global annual wetland dataset at 30 m with a fine classification system from 2000 to 2022. Scientific Data, 11(1): 310
  199. 199.
    Zhao C P, Jia M M, Wang Z M, Mao D H and Wang Y Q. 2023a. Identifying mangroves through knowledge extracted from trained random forest models: an interpretable mangrove mapping approach (IMMA). ISPRS Journal of Photogrammetry and Remote Sensing, 201: 209-225
  200. 200.
    Zhao C P, Qin C Z, Wang Z M, Mao D H, Wang Y Q and Jia M M. 2022. Decision surface optimization in mapping exotic mangrove species (Sonneratia apetala) across latitudinal coastal areas of China. ISPRS Journal of Photogrammetry and Remote Sensing, 193: 269-283
  201. 201.
    Zhao J L, Hu L, Dong Y Y, Huang L S, Weng S Z and Zhang D Y. 2021. A combination method of stacked autoencoder and 3D deep residual network for hyperspectral image classification. International Journal of Applied Earth Observation and Geoinformation, 102: 102459
  202. 202.
    Zhao J L, Hu L, Huang L S, Wang C J and Liang D. 2023b. MSRA-G: combination of multi-scale residual attention network and generative adversarial networks for hyperspectral image classification. Engineering Applications of Artificial Intelligence, 121: 106017
  203. 203.
    Zhao N B, Yang J J, Qin K, Zhao Y J, Yang Y C and Zhu L. 2021. Comprehensive inversion model of Se content in black soil by aerial hyperspectral method. Science of Surveying and Mapping, 46(6): 128-135
  204. 204.
    Zhao W, Bao N S, Liu S J, Mao Y C and Xiao D. 2020. Plant litter effect of the soil organic carbon estimation and unmixing method based on the visible-near infrared spectra. Spectroscopy and Spectral Analysis, 40(7): 2188-2193
  205. 205.
    Zhao Y H, Jia M M, Sun G Y and Zhang A Z. 2025. PAMSNet: a point annotation-driven multi-source network for remote sensing semantic segmentation. ISPRS Journal of Photogrammetry and Remote Sensing, 229: 1-16
  206. 206.
    Zhao Y H, Sun G Y, Ling Z Y, Zhang A Z and Jia X P. 2024. Point-based weakly supervised deep learning for semantic segmentation of remote sensing images. IEEE Transactions on Geoscience and Remote Sensing, 62: 5638416
  207. 207.
    Zheng Z, Zhong Y F, Wang J J, Ma A L and Zhang L P. 2021. Building damage assessment for rapid disaster response with a deep object-based semantic change detection framework: from natural disasters to man-made disasters. Remote Sensing of Environment, 265: 112636
  208. 208.
    Zhou H, Sun M, Ren X and Wang X Y. 2021. Visible-thermal image object detection via the combination of illumination conditions and temperature information. Remote Sensing, 13(18): 3656
  209. 209.
    Zhou M L, Xing R, Han D L, Qi Z Y and Li G. 2025. PDT: UAV target detection dataset for pests and diseases tree//Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T and Varol G, eds. Computer Vision – ECCV 2024. Cham: Springer: 56-72
  210. 210.
    Zhu P F, Peng T, Du D W, Yu H T, Zhang L B and Hu Q H. 2021. Graph regularized flow attention network for video animal counting from drones. IEEE Transactions on Image Processing, 30: 5339-5351
  211. 211.
    Zhu P F, Wen L Y, Du D W, Bian X, Fan H, Hu Q H and Ling H B. 2022a. Detection and tracking meet drones challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11): 7380-7399
  212. 212.
    Zhu Q, Zhang L G, Ding Y L, Hu H, Ge X M, Liu M W and Wang W. 2022. From real 3D modeling to digital twin modeling. Acta Geodaetica et Cartographica Sinica, 51(6): 1040-1049
  213. 213.
    Zhu Q Q, Guo X, Deng W H, Shi S N, Guan Q F, Zhong Y F, Zhang L P and Li D R. 2022c. Land-Use/Land-Cover change detection based on a Siamese global learning framework for high spatial resolution remote sensing imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 184: 63-78
  214. 214.
    Zhu X X, Chen S N, Zhang F H, Shi Y L and Wang Y Y. 2025. GlobalBuildingAtlas: an open global and complete dataset of building polygons, heights and LoD1 3D models. Earth System Science Data, 17(12): 6647-6668
  215. 215.
    Zhu Z and Woodcock C E. 2012. Object-based cloud and cloud shadow detection in Landsat imagery. Remote Sensing of Environment, 118: 83-94
  216. 216.
    Zuo X B, Rui J, Ding L, Jin F, Lin Y Z, Wang S X, Liu X and Lei J. 2025. Integrating segment anything model with instance-level change generation for single-temporal unsupervised change detection. IEEE Transactions on Geoscience and Remote Sensing, 63: 5631717

Lire l'article complet

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website