
人工智能与遥感科学交叉研究进展
Theme Keywords: deep learningremote sensing imagesremote sensing imagerysemantic segmentationhyperspectral imagezero-shot classificationwater remote sensing reflectancewater quality parameter retrievalvisual language modelvision transformers
- The Paper
- Abstract:Perennial soil erosion poses a serious threat to the black soil area of Northeast China, and gully erosion is one of its main manifestations. Remote sensing technology has been widely used in the monitoring and management of gully erosion, and considerable labeled historical survey data have been accumulated. However, how to use these historical data to reliably extract gully information from the latest data captured by various sensors at different times remains an urgent technical problem to be solved. Therefore, this study aims to develop an effective method to achieve reliable cross-temporal gully extraction and provide technical support for land protection and management in black soil areas.On the basis of the above objective, this study proposes a cyclic self-training framework (CSTF). It employs an iterative self-training approach to realize reliable cross-temporal gully extraction. In each self-training iteration, an object-level pseudolabel generation strategy is designed to ensure high-quality pseudolabels for the latest data. Additionally, a loss function based on pseudolabel confidence factors is introduced to effectively mitigate the adverse effects of pseudolabel noise.In the experimental section, Huachuan County, Heilongjiang Province, China was selected as the study area, and the following results were drawn: (1) The characteristic differences between historical and current data present significant challenges for cross-temporal extraction tasks. Compared with traditional fully supervised methods, unsupervised domain adaptation methods offer superior performance. Moreover, the self-training methods demonstrate greater robustness than invariant representation learning, thus justifying their use in cross-temporal extraction studies. (2) For self-training methods, the quality of pseudolabels is a critical factor influencing performance. Hence, a series of improvement strategies is proposed, leading to the best results in accuracy assessments and visual interpretation. Specifically, the intersection over union is 7.39% and 7.90% higher than those of the second-best methods in Experiments 1 and 2, respectively. Furthermore, these strategies are shown to be effective, necessary, and compatible, as demonstrated through detailed ablation experiments. Regarding the analysis of algorithm complexity and operational efficiency, the proposed CSTF cannot only ensure accurate extraction results but can also offer high efficiency, meeting the actual monitoring requirements for gully erosion.In conclusion, the proposed CSTF provides robust technical support for cultivated land conservation in black soil regions and offers a promising approach for sustainable land management. Currently, CSTF only deals with the binary classification of erosion gullies and noneroded areas. Future research will expand the framework to recognize gullies at different developmental stages, facilitating refined monitoring and analysis of erosion gullies.Keywords:Soil erosion;black soil area of Northeast China;gully erosion;cross-temporal extraction;self-training;object-level pseudo-label generation strategy;pseudo-label credibility factor;pseudo-label noise300|458|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Estimating building height using optical and SAR remote sensing imagery is of great significance for understanding urban morphology and optimizing urban space utilization. However, current datasets often suffer from small sample sizes, limited geographic diversity, and a lack of openness, making them insufficient for supporting deep learning-based remote sensing applications—especially for large-scale studies in China. Accurately estimating building heights is critical for understanding urban morphology and optimizing urban stock space, thus necessitating the development of a comprehensive, representative, and accessible dataset. Methods To overcome these issues, this study constructs a building height dataset based on Sentinel imagery for deep learning (Building Height Estimation Dataset Based on Sentinel Imagery, BHDSI), specifically designed for building height regression tasks. The dataset comprises 5,606 samples from the central urban areas of 62 cities across China, making it the largest building height dataset for the country in terms of geographic coverage. It includes Sentinel-1 and -2 imagery along with true building height values, with each sample having a spatial resolution of 256×256 pixels. This dataset provides an important supplementary choice compared with existing datasets with smaller sample sizes, such as 64×64. The dataset encompasses a wide range of scenarios, including urban and rural areas, ensuring effective representation of spatial features. Results Experimental evaluations demonstrate that the BHDSI dataset leads to superior performance in building height regression tasks in comparison with other similar datasets across various deep learning networks. The results indicate that estimation accuracy tends to be high in regions with low building heights. Furthermore, this study determines that using a U-Net decoder structure in the network architecture contributes to enhanced prediction precision, highlighting the importance of decoder design in deep learning-based height estimation. Conclusion The BHDSI dataset significantly advances the field of building height estimation by offering a large-scale, diverse, and high-quality resource tailored for deep learning. Its broad coverage, balanced height distribution, large sample size, and open accessibility make it better suited for training and evaluating deep neural networks than previously available datasets. This study confirms that data quality and network architecture, especially decoder design, play vital roles in improving estimation accuracy, and BHDSI serves as a strong foundation for future research in this domain.Keywords:Sentinel imagery;building height;dataset;deep learning;convolutional neural network1065|1039|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Land cover mapping is a vital task in Earth observation; it provides fine-scale details of landscape and support many downstream applications like ecology, hydrology, and resource management. However, current land cover mapping faces key challenges such as limited information from single-source data, substantial data heterogeneity, and insufficient generalization capability of individual models.To solve these issues, this study addresses key challenges in land cover mapping by proposing an innovative framework that integrates multimodal remote sensing data with a multimodel deep learning-based framework for collaborative decision-making. This work aims to provide a novel pathway for large-scale high-resolution land cover mapping and support many other downstream applications.In this study, leveraging the spectral characteristics of multispectral imagery (MSI) and the distinct properties of Synthetic Aperture Radar (SAR) data, a complementary multimodal (i.e., MSI+SAR) dataset is constructed as feature input, effectively overcoming the limitations of using a single SAR modality in complex Earth observation scenarios. Moreover, at the model architecture level, a systematic evaluation of the performance differences among seven representative machine learning-based models is conducted. Based on these foundations, a multimodel fusion strategy is further proposed, combining convolutional neural networks (CNNs), vision transformers (ViTs), and a hybrid CNN/ViT architecture (represented by FCN, ConViT, and CoAtNet, respectively). These three models have demonstrated outstanding performance in previous comparative experiments between multimodal (i.e., SAR+MSI) data and single-modal SAR data.In the experimental section, we conduct a comprehensive evaluation in the entire Beijing City. Remote sensing data are collected from GF-3 (SAR) and GF-6 (MSI). On the basis of these multimodal data, we compare six advanced semantic segmentation models and one pixel-based classification model. Results show that compared with using single SAR data, the utilization of multimodal data significantly improves two key evaluation metrics used in semantic segmentation task or land cover mapping task—Overall Accuracy (OA) and Frequency-Weighted Intersection over Union (FWIoU)—in Beijing City. The proposed multimodel fusion framework further enhances performance in OA and FWIoU, validating the effectiveness of the method in semantic segmentation tasks with multimodal remote sensing data.This paper presents an innovative approach that enhances a model’s capability to extract complex land cover features by effectively integrating multimodal remote sensing data (MSI and SAR) with a multimodel fusion framework. The proposed method demonstrates superior performance in large-scale land cover mapping, achieving significant improvements in classification accuracy and robustness compared with conventional single-source or -model approaches. The success of this framework highlights the powerful potential of multimodal data fusion and collaborative deep learning in overcoming challenges such as spectral ambiguity, cloud interference, and limited labeled samples.Keywords:land-cover mapping;multi-spectral;synthetic aperture radar;multi-modal;multi-model fusion;convolutional neural networks;vision transformers796|927|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Deep learning-based object detection has become an important tool for large-scale burial site identification in remote sensing archaeology, with its performance highly dependent on sufficient and diverse annotated datasets. However, in practical archaeological applications, acquiring large-scale, high-quality samples is both expensive and time-consuming. Moreover, burial sites are often distributed across highly heterogeneous environments, leading to significant environmental imbalance within datasets. Under few-shot conditions, models tend to overfit to dominant background features and exhibit limited cross-scene generalization capabilities. This study aims to develop an environment-semantic data augmentation strategy that expands background diversity while preserving the original spatial structure and label distribution of burial targets. By simulating multiple environmental contexts through generative modeling, the proposed method can alleviate sample imbalance to a certain extent, enhance robustness to environmental variations, and improve cross-domain transferability, thereby providing a practical and scalable solution for few-shot remote sensing archaeological detection.This study proposes a diffusion-based environment-semantic augmentation framework consisting of three components: environmental generation, fractal fusion, and random enhancement. A pre-trained InstructPix2Pix diffusion model is guided by predefined environmental prompts to simulate burial sites under diverse background conditions while preserving structural features. To further enhance robustness, fractal patterns are fused with images using a Beta-distribution-based weighted multiplication strategy, allowing texture-level enhancement without occluding targets. Additional random image operations are applied to increase variability.The method was evaluated on a self-constructed Altai burial dataset using multiple object detection models, with Mosaic, MixUp, and DiffuseMix serving as baselines. A WorldView-2 dataset was used for cross-domain testing.The proposed method consistently improved detection performance across various models. Compared with Mosaic, the average AP50 across the model ensemble increased by 7.4% on the test set and 12.2% on the validation set, with AP50-95 achieving a maximum improvement of 19.1%. In transfer tasks on heterogeneous datasets, AP50 improved by 16.4%, outperforming both MixUp and DiffuseMix.Through environment-semantic augmentation, the method mitigated background variations across diverse natural environments and enhanced the model’s capability to discern burial targets. Transfer experiments on WorldView-2 imagery further demonstrated that the approach improves model generalization and prevents excessive focus on background features.This study presents an environment-semantic augmentation framework that integrates diffusion-based background simulation with fractal texture fusion for remote sensing archaeological detection. By enriching environmental diversity while preserving target structure, the proposed approach effectively addresses few-shot learning, environmental imbalance, and overfitting issues that commonly limit archaeological object detection models.Experimental results across multiple detection architectures and heterogeneous datasets demonstrate significant improvements in accuracy, recall, robustness, and cross-dataset generalization. The findings confirm that generative models, when properly guided by environmental semantics, can provide meaningful and controllable data diversity for small-sample scenarios.Keywords:deep learning;ancient tomb detection;data augmentation;diffusion model;remote sensing archaeology;few-shot learning;environmental semantic enhancement418|742|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:To address the challenges in acquiring hyperspectral remote sensing reflectance (Rrs) data and the limitations of existing spectral reconstruction methods, including their reliance on in-situ measurements, weak generalization ability, and insufficient accuracy in optically complex coastal waters, this study proposes a novel deep learning model based on the Kolmogorov-Arnold Network (KAN). The model is designed to directly exploit widely available multispectral satellite observations to efficiently reconstruct continuous hyperspectral Rrs that closely match the characteristics of true measurements, thereby overcoming the bottlenecks of traditional approaches and improving remote sensing inversion performance in complex nearshore environments.The proposed method adopts an end-to-end nonlinear modeling framework within the KAN architecture, incorporating learnable nonlinear activation functions to flexibly capture complex local nonlinear relationships in the input data. This approach enables accurate reconstruction of continuous hyperspectral Rrs from multispectral inputs, with spectral distributions highly consistent with actual observations. In this work, Level-2 Rrs products from the Hyperspectral Imager for the Coastal Ocean were used as training samples. These products were resampled in accordance with the spectral response functions of six mainstream multispectral sensors (S3A OLCI, MERIS, MODIS, SeaWiFS, S2A MSI, and OLI) to generate “multispectral-hyperspectral” data pairs for model training. The approach eliminates the need for in-situ or optical simulation data, relying solely on remote sensing observations for training and modeling, thus greatly enhancing model generality and practicality.Experimental results demonstrate that the KAN model achieved superior reconstruction accuracy (mean R2 > 0.9982) and robustness across all six sensors compared with benchmark models. It not only reproduced the overall spectral shape with high fidelity but also exhibited strong detail-capturing capability in critical regions with missing or sparse sensor bands, such as the red-edge region beyond 680 nm, producing reconstructed curves nearly identical to the original spectra. In downstream chlorophyll-a inversion applications, the use of KAN-reconstructed data significantly improved retrieval accuracy over original multispectral inputs, reducing RMSE by approximately 16.13% and increasing R2 by 3.30%. The advantages were particularly evident in high-concentration waters, effectively overcoming the limitations of multispectral sensors.Overall, the proposed KAN-based hyperspectral Rrs reconstruction model successfully breaks the dependency bottleneck of traditional methods on in-situ or simulated data, offering a highly accurate and generalizable universal solution. By generating high-quality continuous spectra from existing multispectral datasets, it substantially enhances the performance and accuracy of water quality parameter retrieval in complex aquatic environments. The model serves as a powerful data augmentation tool for aquatic remote sensing and a novel technical pathway for leveraging vast archives of historical multispectral satellite data to enable reliable global water environment monitoring.Keywords:KAN network;hyperspectral remote sensing reflectance;remote sensing reflectance reconstruction;water quality parameter retrieval;coastal water bodies;water remote sensing reflectance387|623|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Automatic road extraction from high-resolution remote sensing images plays a crucial role in applications such as smart cities, intelligent transportation, and autonomous driving. However, existing methods often suffer from issues like fragmentation and poor connectivity in the extracted road networks, especially under complex scenarios with occlusions, shadows, and large-scale variations. This study aims to develop a robust deep learning model capable of extracting continuous and complete road networks from high-resolution remote sensing imagery by effectively integrating multiscale contextual information and attention mechanisms.An improved encoder-decoder network named Split-Attention and Multi-Scale Attention Network (SAMSNet) is proposed. The encoder is based on ResNeSt-50, which utilizes a split-attention mechanism to enhance cross-channel feature interaction and capture rich semantic representations. A cascaded parallel dilated convolution block (Dblock) is introduced in the central part of the network to expand the receptive field and aggregate multiscale context without losing spatial details. Furthermore, a multiscale channel attention module (MS-CAM) is incorporated into the skip connections to simultaneously emphasize global and local road features, improving the model’s ability to handle extreme scale variations. The network is trained using a combined loss function of binary cross-entropy and Dice loss to address class imbalance and emphasize boundary accuracy.Extensive experiments were conducted on three public road extraction datasets DeepGlobe, Massachusetts, and GRSet. SAMSNet achieved state-of-the-art performance across all datasets. On the DeepGlobe dataset, it attained an IoU of 74.48% and an F1-score of 85.37%, significantly outperforming other models such as U-Net, D-LinkNet, and transformer-based approaches. Similar improvements were observed on the Massachusetts dataset, with IoU and F1-score reaching 66.61% and 79.96%, respectively. Transfer learning experiments on the GRSet dataset further demonstrated the strong generalization capability of SAMSNet, where it achieved the highest IoU (55.55%) and F1-score (60.71%) among all compared models. Ablation studies confirmed the individual contributions of the Dblock and MS-CAM modules to the overall performance.SAMSNet effectively integrates split-attention, multiscale dilated convolution, and channel attention mechanisms to improve the accuracy, connectivity, and completeness of road extraction from high-resolution remote sensing images. The proposed model shows strong performance across diverse datasets and complex scenarios, indicating its robustness and generalization ability. However, the model’s high computational complexity may limit its deployment in real-time applications. Future work will focus on developing lighter versions of the model and exploring joint extraction of road segmentation and centerline detection.Keywords:remote sensing images;road extraction;semantic segmentation;ResNeSt-50;Dispersed Attention;Multi-Scale Channel Attention;dilated convolution434|579|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Multitemporal hyperspectral images, characterized by their rich spectral information and high spatial resolution, have demonstrated significant potential for change detection tasks across diverse environmental and urban monitoring scenarios. Traditional hyperspectral change detection (HSICD) algorithms based on supervised learning often rely on a large number of labeled samples, which requires a large sample annotation cost. Although recent studies have begun to address change detection under limited labeled samples, existing approaches still fail to fully exploit the limited supervisory signals and often lack effective mechanisms for capturing discriminative change-related features. To address these challenges, this paper plans to design a novel network architecture that can maximize the utility of limited labeled samples while enhancing the representation of differential features between bi-temporal hyperspectral images, thereby achieving high-precision change detection under constrained annotation conditions.In this paper, we propose a joint central difference feature and spatial-spectral attention network (JCDS2AN) for HSICD, which can alleviate the fluctuation in changing features under sample constraints and learn representative changing features. In JCDS2AN, a multiscale spatial-spectral attention block is designed to simultaneously capture fine-grained spatial details and continuous spectral signatures at varying scales, enabling the model to learn robust and discriminative feature representations from scarce training data. Second, to amplify change-related information while suppressing background interference, a differential feature enhancement module is introduced, which explicitly models the temporal spectral variations. Third, and most distinctively, a differential center pixel exchange strategy is proposed. This strategy leverages the extracted differential features to guide the information exchange between bi-temporal feature representations, effectively aligning and reinforcing change-salient regions. By integrating these designs, JCDS2AN forms an end-to-end trainable framework that holistically addresses the dual challenges of sample scarcity and feature ambiguity in hyperspectral change detection. Experimental results on three publicly available hyperspectral image datasets show that the proposed JCDS2AN outperforms the state-of-the-art methods in HSICD. When utilizing only 1% of the training samples, the method achieved optimal Kappa and OA of 95.90% and 98.30%, respectively, on the Farmland dataset. Ablation experiments were conducted for each proposed module to demonstrate their effectiveness. This approach is capable of extracting discriminative deep change semantic information, with both qualitative and quantitative results surpassing those of other advanced networks. This paper proposes JCDS2AN, specifically designed to operate effectively under the constraint of limited labeled samples. The network fully exploits the differential information between bi-temporal images, enabling effective information interaction between bi-temporal features and differential features. It learns representative change features from paired input HSIs and alleviates intra-class feature fluctuations caused by limited training samples. Validation experiments are conducted on three public datasets, including comparative analysis with state-of-the-art change detection methods and ablation studies.Keywords:hyperspectral image;remote sensing images;change detection;Multi-scale Features;differential feature guidance;center pixel470|701|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Remote sensing image interpretation plays a vital role in Earth observation, yet its performance is frequently hampered by complex meteorological conditions. Under foggy weather, atmospheric scattering significantly weakens the illumination intensity and reduces the contrast of images. This degradation leading to the loss of critical target details and color distortion, which poses a significant challenge to the performance of object detection models. Targets such as small vehicles or ships often become indistinguishable from the background noise, leading to high miss rates and false alarms in practical applications.Current research primarily addresses this issue through two strategiestraining detection models directly on foggy datasets to enhance robustness, or utilizing image dehazing as a preprocessing step to recover visibility. However, these approaches have inherent limitations. Directly training on degraded data often fails to capture fine-grained features, while the dehazing process can result in loss of feature due to the decoupling of restoration and detection. Specifically, it is difficult to ensure a consistently positive correlation between dehazing results and object detection tasks, i.e., the dehazing results are not always beneficial for object detection because traditional dehazing focuses on human visual perception rather than machine-learning-oriented semantic clarity.To address this issue, this study proposes a cascade learning foggy object detection method (CL-FODM). This method establishes an end-to-end framework that integrates a lightweight dehazing subnetwork with a high-performance detection subnetwork. The dehazing subnetwork combines CNN and Transformer architectures, specifically incorporating the MB-TaylorFormer module. By leveraging the Taylor expansion to approximate the self-attention mechanism, the model effectively captures global dependencies with reduced computational complexity, which can obtain clear dehazed features and provide salient semantic information for the object detection task. For the downstream detection, we employ DiffusionDet, which treats object detection as a generative denoising process, offering superior flexibility and accuracy in complex scenes. Furthermore, a multitask loss function guided by feature perception is constructed to precisely mine discriminative target semantic features at the feature level. By introducing a shared feature perception module at the encoder stage, the framework calculates the perceptual difference between the restored features and the ground-truth clear features. This mechanism achieving collaborative optimization between dehazing and object detection and solving the semantic inconsistency between low- and high-level tasks. It ensures that the dehazing subnetwork is optimized toward a direction that maximizes detection accuracy rather than just pixel-level similarity.Experimental results, conducted on DOTA1.0 with various fog densities, show that the proposed CL-FODM outperforms the original model and the cascaded model in terms of evaluation metrics and visual detection effects. On the thick fog dataset, the CL-FODM achieved a mAP of 40.291, which is a 1.327 improvement over the baseline DiffusionDet, while adding only 0.217 M parameters. These results demonstrate that the proposed CL-FODM effectively restores target visibility and preserves discriminative semantic information, providing a robust solution for remote sensing object detection in adverse weather conditions.Keywords:remote sensing imagery;object detection;Dehazing Model;deep learning;Cascade Learning386|622|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Semantic segmentation of Remote Sensing Images (RSIs) constitutes a fundamental yet highly challenging computer vision task, serving as the technological backbone for numerous critical applications including fine-grained land cover and land use classification, large-scale urban planning, infrastructure monitoring, and multi-temporal change detection. In recent years, unsupervised domain adaptation has emerged as a promising paradigm to alleviate the heavy reliance on pixel-level manual annotations by transferring knowledge from labeled source domains to unlabeled target domains. Despite these advances, existing domain adaptation methods predominantly adopt single-task learning architectures, which inherently suffer from limited feature representation capacity. This limitation becomes particularly pronounced when dealing with hard-to-classify regions in RSIs—such as densely built urban areas with complex textures, regions heavily affected by cast shadows, vegetation occlusions, ambiguous object boundaries, and spectrally similar land cover categories. The absence of complementary information in single-task models severely constrains their ability to disambiguate these challenging cases. To fundamentally address this issue, this study proposes a novel multi-task learning domain adaptive network (MTLDANet), which jointly learns and exploits semantic and elevation information from RSIs. By leveraging the intrinsic correlation between these two complementary tasks, the proposed framework significantly enhances feature discriminability, improves segmentation robustness, and achieves superior cross-domain generalization under various imaging conditions. The proposed MTLDANet adopts a shared encoder architecture with two task-specific decoders dedicated to semantic segmentation and elevation estimation, respectively. The extracted task-specific features are first processed independently, then jointly fed into a carefully designed cross-task feature correlation learning module. This module explicitly models the latent correlations between semantic categories and elevation patterns, enabling bidirectional knowledge transfer and mutual reinforcement of feature representations. Specifically, elevation cues provide geometric priors that help distinguish objects with similar spectral signatures but different vertical structures, while semantic context aids in resolving ambiguities in elevation prediction. Furthermore, to address the domain shift problem, a hybrid consistency learning module guided by high-confidence pseudo-labels is introduced. This module operates at both the feature level and output space, iteratively refining pseudo-label quality through self-training and enforcing global domain alignment via adversarial learning. In addition, recognizing that certain land cover categories (e.g., roads and building shadows, low vegetation and trees) are persistently confused during domain adaptation, an entropy-guided category-level alignment module is incorporated. This module computes prediction entropy to identify regions of high uncertainty, and performs fine-grained category-wise distribution alignment between source and target domains. By focusing alignment efforts on ambiguous categories, the model significantly enhances feature separability and classification reliability in target domains. To comprehensively validate the effectiveness and generalization ability of the proposed method, extensive experiments are conducted on four challenging cross-scene RSI segmentation tasks using the ISPRS 2D semantic labeling dataset and the US3D dataset. These experiments cover a wide range of urban and suburban scenes captured under different sensors, geographical locations, lighting conditions, and seasonal variations. Quantitative evaluations using standard metrics such as mean intersection over union (mIoU) and F1-score demonstrate that MTLDANet consistently and substantially outperforms existing state-of-the-art unsupervised domain adaptation approaches across all experimental settings. Ablation studies further confirm the individual contribution of each proposed module. Both quantitative and qualitative analysis of the experimental results provide strong evidence that MTLDANet achieves significant and consistent advantages in diverse complex cross-domain scenarios. The integration of multi-task learning with cross-task feature correlation, hybrid consistency regularization, and entropy-guided category-level alignment forms a highly effective framework for advancing domain adaptive remote sensing image segmentation. This work not only sets a new state-of-the-art benchmark but also offers valuable insights for future research on multi-modal and multi-task learning in remote sensing.Keywords:semantic segmentation;unsupervised domain adaptation;remote sensing imagery;Multi-task learning;elevation information;semantic information;pseudo-label;Entropy583|418|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23 - Abstract:Remote Sensing Image Referring Segmentation (RRSIS) aims to accurately locate and delineate specific regions within high-resolution remote sensing imagery on the basis of natural language referring expressions and ultimately achieve pixel-level semantic interpretation. This task critically bridges user demands and intelligent geospatial information analysis. However, compared with natural scene referential segmentation, RRSIS presents two unique challenges.(1) Relatively low contrast between targets and their surroundings often leads to a semantic dispersion phenomenon, where the segmentation mask covers irrelevant areas.(2) Substantial cross-modal semantic gaps exist between visual and textual representations. Conventional cross-modal attention mechanisms tend to rely on coarse feature alignments, which are insufficient for fine-grained geographical boundary delineation.The objective of this study is to design a robust and generalizable framework that can effectively mitigate semantic dispersion, narrow the modality gap, and achieve precise alignment between entity-level textual descriptions and complex geospatial visual features in RRSIS tasks.The proposed Enti-CroM, an entity-guided cross-modal interaction framework tailored for RRSIS, is adopted in this study.Entity-Guided Self-Reasoning (SEG) module: Motivated by the Segment Anything Model (SAM), the SEG module injects fine-grained entity priors into the model by leveraging spatial-structural constraints. A self-reasoning process generates robust and coherent entity prompts, which are integrated with visual and textual embeddings to form a trimodal entity–vision–text feature cube. Hierarchical Modality Interaction (HMI) mechanism: Parameter-Free Mutual Activation (PFMA): PFMA is a neuroscience-inspired and spatially aware mutual modulation approach that computes positionwise semantic similarity between modalities without introducing additional learnable parameters. PFMA enables efficient and precise semantic information propagation, suppresses irrelevant background interference, and reduces modality misalignment. Entity-Guided Cross-Attention (EGCA): EGCA incorporates the entity prior as an attention guide to refine the interaction between textual and visual streams and ultimately enhance the ability of the model to represent irregular and fine-grained geographical boundaries. The overall architecture decouples cross-modal semantic propagation from fine-grained spatial dependency modeling to ensure high-level semantic consistency and spatial precision.Extensive experiments were conducted on two benchmark datasets, namely, RefSegRS and RRSIS-D, which are widely used for RRSIS evaluation. Performance was assessed via the mean intersection-over-union (mIoU) metric. Compared with the strongest existing state-of-the-art method, Enti-CroM achieved absolute mIoU improvements of +3.23% on RefSegRS and +2.62% on RRSIS-D. Ablation studies further confirmed the effectiveness of each component. The SEG module alone significantly improved target localization and robustness to background clutter. The HMI mechanism, particularly PFMA, improved modality alignment and suppression of semantic noise, whereas EGCA improved boundary representation in complex spatial contexts. Qualitative visual comparisons demonstrated that Enti-CroM delivers sharper object boundaries, more accurate correspondence to the referring expressions, and fewer false positive regions, especially in heterogeneous landscapes such as urban areas and agricultural mosaics.This work addresses two longstanding challenges in RRSIS, namely, semantic dispersion and cross-modal gaps, by integrating entity-guided priors and a hierarchical modality interaction strategy. Incorporating spatially grounded entity cues and explicit, fine-grained semantic alignment allows Enti-CroM to substantially enhance segmentation accuracy and robustness in complex remote sensing scenes. The proposed framework not only sets new benchmarks on two challenging datasets but also offers a general paradigm for entity-aware multimodal analysis in remote sensing. Despite the advantages of the Enti-CroM, it still faces certain limitations, such as reliance on the quality of entity priors and increased computational demand for ultrahigh-resolution imagery. Future work will focus on three aspects: (1) developing adaptive or self-supervised entity prior generation mechanisms to reduce dependency on external annotations; (2) incorporating model compression and acceleration for large-scale deployment; and (3) extending the framework to integrate additional modalities, such as hyperspectral and SAR data, and broaden earth observation applications.Keywords:remote sensing imagery;referring segmentation;cross-modal interaction;SAM;entity awareness;attention mechanisms;spatial-structural constraints527|870|0
<HTML><L-PDF><Enhanced-PDF><Meta-XML>Updated:2026-03-23



