DSP-UNet:A Remotely-sensed Water Body Extraction Model Based on U-Net Enhanced with Dual Attention and Strip Pooling

HU Jin-peng, YE Song, ZHENG Xue-dong, LI Qin, WANG Ying, ZHANG Wei, PAN Zhi-quan

Journal of Changjiang River Scientific Research Institute ›› 2026, Vol. 43 ›› Issue (9) : 158-169.

PDF(3114 KB)
PDF(3114 KB)
Journal of Changjiang River Scientific Research Institute ›› 2026, Vol. 43 ›› Issue (9) : 158-169. DOI: 10.11988/ckyyb.20250650
Water Conservancy Informatization

DSP-UNet:A Remotely-sensed Water Body Extraction Model Based on U-Net Enhanced with Dual Attention and Strip Pooling

Author information +
History +

Abstract

[Objective] Although deep learning-based semantic segmentation models, such as U-Net and DeepLabv3+, have improved water body extraction performance, they remain limited in delineating continuous boundaries and recognizing narrow water bodies. To address these limitations, this study proposes an improved model named DSP-UNet (U-Net enhanced with Dual Attention and Strip Pooling). The model is designed to improve both local boundary sensitivity and global contextual reasoning, achieving higher precision and robustness while maintaining computational efficiency. The objective is to develop a high-performance deep learning model capable of stable water body extraction under varying geographic and imaging conditions. [Methods] The proposed DSP-UNet model was constructed based on the classical U-Net architecture and integrated three key innovations: strip pooling module (SPM), dual attention mechanism (DAM), and SimAM attention module. The model was trained using the OpenWUSU512 dataset. Comparative experiments with baseline models, including U-Net, PSPNet, HRNet, and DeepLabv3+, were performed under the same training conditions. Ablation studies were conducted to quantify the contribution of each proposed module. Moreover, cross-domain generalization was evaluated on the Satellite Images of Water Bodies dataset to assess model transferability. [Results] Experimental results demonstrated that DSP-UNet achieved higher segmentation accuracy and robustness across all evaluation indicators. On the OpenWUSU512 dataset, DSP-UNet reached a mean intersection over union (MIoU) of 0.962 and an F1 score of 0.966, outperforming PSPNet by 2.4% in MIoU. The precision and recall values were 0.976 and 0.956, respectively, indicating a balanced trade-off between false positives and false negatives. The overall accuracy (OA) was 0.991. Compared with the baseline U-Net, the dual attention mechanism increased MIoU by 1.1%; the integration of SimAM improved edge recognition and reduced boundary noise; and the addition of strip pooling expanded spatial perception, increased MIoU by 2.3%, and reduced total training time by 11.5%. These results confirmed that the progressive introduction of modules effectively enhanced both feature representation capability and computational efficiency. Cross-domain experiments on the Satellite Images of Water Bodies dataset demonstrated strong generalization performance. DSP-UNet achieved an F1 score of 0.909, an IoU of 0.833, and an OA of 0.954, outperforming U-Net and SegNet by more than 5% in IoU. Visual assessment showed that the model accurately distinguished water bodies from shadows and reflections, maintained continuous boundaries, and effectively extracted narrow channels and small ponds, which were often misclassified by conventional models. [Conclusion] The proposed DSP-UNet provides an accurate and robust solution for water body extraction from high-resolution remote sensing imagery. By combining dual attention and strip pooling, the model effectively enhances both global contextual understanding and local boundary refinement. The parameter-free SimAM module further improves spatial discrimination without introducing additional computational burden. DSP-UNet achieves the highest accuracy among the compared models while maintaining stable convergence and reduced training time, demonstrating its efficiency for large-scale applications. The model’s strong generalization ability across datasets indicates its potential application in remote sensing tasks, including wetland mapping, flood detection, and shoreline monitoring. Future research should focus on integrating multi-temporal and multi-sensor data, developing temporal attention mechanisms for dynamic water monitoring, and optimizing the architecture for lightweight and real-time deployment in practical remote sensing applications.

Key words

water body extraction / dual attention mechanism / strip pooling / remote sensing imagery / DSP-UNet

Cite this article

Download Citations
HU Jin-peng , YE Song , ZHENG Xue-dong , et al . DSP-UNet:A Remotely-sensed Water Body Extraction Model Based on U-Net Enhanced with Dual Attention and Strip Pooling[J]. Journal of Changjiang River Scientific Research Institute. 2026, 43(9): 158-169 https://doi.org/10.11988/ckyyb.20250650

References

[1]
Chung M G, Frank K A, Pokhrel Y, et al. Natural infrastructure in sustaining global urban freshwater ecosystem services[J]. Nature Sustainability, 2021, 4(12): 1068-1075.
[2]
刘兆孝. 新时期长江流域水资源保护规划及管理工作的思考[J]. 长江科学院院报, 2024, 41(4): 1-7.
Abstract
组织编制并实施水资源保护规划是国家赋予水行政主管部门的法定职责,也是水利部门履行水资源管理和生态环境保护责任的重要工作。长江流域水资源保护规划经历了从起步、探索、实践与发展到形成完善体系的过程,推进了流域水资源保护与管理科学有序发展。随着国家机构改革部门职能调整和《中华人民共和国长江保护法》出台,为更好地推进和落实新时期长江流域水资源保护规划和管理工作,系统梳理长江流域水资源保护规划发展历程与成效,分析当前流域水资源保护面临的形势与挑战,厘清流域水资源保护工作定位和思路,并据此提出新时期流域水资源保护规划及管理的工作建议,对凝聚流域水资源保护规划共识、建立流域水资源保护规划体系、推进水资源保护高质量发展、助力实现治江现代化具有重要的指导意义。
(Liu Zhaoxiao. Reflections on planning and management of water resources protection in the Yangtze River Basin in the new era[J]. Journal of Changjiang River Scientific Research Institute, 2024, 41(4): 1-7. (in Chinese))
[3]
许继军, 陈述. 新时期长江流域水资源保护利用管理体制机制研究[J]. 长江科学院院报, 2022, 39(7): 1-6.
Abstract
针对当前长江流域水资源保护利用管理存在的问题和诸多难题,对新时期长江流域水资源管理体制机制重构和完善路径进行了探讨。为破解制约流域水资源保护利用可持续发展的体制机制障碍,长江流域水资源管理应遵循流域整体性和复杂性规律,在流域协调机制框架上,健全统分结合、整体联动的管理体制,打造多元参与的流域治理共同体,健全完善多元化生态补偿机制、跨省河湖长制监督机制、控制性水工程统一调度管理机制、流域水资源保护利用空间管控及市场运行机制,从而推动长江流域治理体系现代化。
(Xu Jijun, Chen Shu. Management system and mechanism of water resources protection and utilization in the Yangtze River Basin in the new era[J]. Journal of Yangtze River Scientific Research Institute, 2022, 39(7): 1-6. (in Chinese))
Protecting and utilizing the water resources in Yangtze River Basin has seen many problems and difficulties. In view of this,we discussed the paths of reconstructing and improving the management system and mechanism in the new era. To break the system and mechanism barriers in sustainable development,the management of water resources in the Yangtze River Basin should conform to the integrity and complexity of the basin. In the framework of coordination across the basin,we recommend to 1)build a sound system which integrates unified and separate management with overall linkage and interactions to forge a governance community for the basin with multiple stakeholders;2)enhance the mechanisms of diversified ecological compensation,cross-provincial river-lake-chief supervision,unified dispatch and management of controlling water projects,as well as spatial control and market operation. With such endeavors,the governance of water resources in Yangtze River Basin is expected to be further modernized.
[4]
李欢, 万玮, 冀锐, 等. 中国卫星遥感地表水资源监测能力分析与展望[J]. 遥感学报, 2023, 27(7):1554-1573.
(Li Huan, Wan Wei, Ji Rui, et al. Inspects and prospects of satellite remote sensing monitoring ability for land surface water in China[J]. National Remote Sensing Bulletin, 2023, 27(7): 1554-1573. (in Chinese))
[5]
许全喜, 许继军. 发展长江水利新质生产力的几点思考[J]. 长江科学院院报, 2024, 41(9): 1-7.
Abstract
发展水利新质生产力,是破解长江新老水问题、提升流域安全保障能力、推动长江经济带高质量发展的内在要求和重要着力点。结合长江治理保护新阶段特征和面临的挑战,基于新质生产力概念,从生态水利科技创新、生产要素数字智能化、产业绿色转型升级等方面,探讨长江水利新质生产力的形成路径和原则要求,并提出加强近自然水工程、数字孪生、节水减污降碳、流域现代管理等方面的科技创新建议,以促进长江生态化、智慧化、绿色化和现代化发展。
(Xu Quanxi, Xu Jijun. Insights into developing new productive forces of water conservancyof the Yangtze River[J]. Journal of Changjiang River Scientific Research Institute, 2024, 41(9): 1-7. (in Chinese))
[6]
Speight L J, Cranston M D, White C J, et al. Operational and emerging capabilities for surface water flood forecasting[J]. WIREs Water, 2021, 8(3): e1517.
[7]
张宝文, 赵展. 利用多地物指数自动提取水体方法[J]. 测绘通报, 2024(11): 21-26.
Abstract
目前常用的水体提取仍然是半自动的方式,其效率较低且容易出现漏提与误提的现象。本文提出一种利用多地物指数自动提取水体的方法,利用植被指数、水体指数、改进的水体指数及建筑用地指数进行水体与非水体样本的自动提取,训练分类器,进而实现高精度全自动的城市水体提取。选择国内外4种典型的不同地理条件的试验区域,与传统的提取方法进行对比分析,该方法下4个试验区域均取得了较高的提取精度,Kappa系数均达0.94以上。与单一波段阈值法相比,Kappa系数平均提升7.2%;与水体指数法相比,Kappa系数平均提升3.8%。
(Zhang Baowen, Zhao Zhan. Automatic extraction of water body using multi-feature index[J]. Bulletin of Surveying and Mapping, 2024(11): 21-26. (in Chinese))
At present, water extraction is still a semi-automatic method, which has low efficiency and is prone to omissions and extraction errors. In this paper, an automatic water extraction method using multi-feature index is proposed. Vegetation index, water index, improved water index and building land index are used to automatically extract water and non-water samples, and the classifier is trained to realize high-precision and automatic urban water extraction. Four typical experimental areas with different geographical conditions at home and abroad are selected for comparison and analysis with the traditional extraction method. Under this method, high extraction accuracy is achieved in all the four experimental areas, and Kappa coefficient reached above 0.94. Compared with single-band threshold method, Kappa coefficient increases by 7.2% on average. Compared with the water index method, Kappa coefficient increased by 3.8% on average.
[8]
Laonamsai J, Julphunthong P, Saprathet T, et al. Utilizing NDWI,MNDWI,SAVI,WRI,and AWEI for estimating erosion and deposition in Ping River in Thailand[J]. Hydrology, 2023, 10(3):70.
The Ping River, located in northern Thailand, is facing various challenges due to the impacts of climate change, dam operations, and sand mining, leading to riverbank erosion and deposition. To monitor the riverbank erosion and accretion, this study employs remote sensing and GIS technology, utilizing five water indices: the Normalized Difference Water Index (NDWI), Modified Normalized Difference Water Index (MNDWI), Soil-Adjusted Vegetation Index (SAVI), Water Ratio Index (WRI), and Automated Water Extraction Index (AWEI). The results from each water index were comparable, with an accuracy ranging from 79.10 to 94.53 percent and analytical precision between 96.05 and 100 percent. The AWEI and WRI streams showed the highest precision out of the five indices due to their larger total surface water area. Between 2015 and 2022, the riverbank of the Ping River saw 5.18 km2 of erosion. Conversely, the morphological analysis revealed 5.55 km2 of accretion in low-lying river areas. The presence of riverbank stabilizing structures has resulted in accretion being greater than erosion, leading to the formation of riverbars along the Ping River. The presence of water hyacinth, narrow river width, and different water levels between the given periods may impact the accuracy of retrieved river areas.
[9]
Gao Bocai. NDWI:A normalized difference water index for remote sensing of vegetation liquid water from space[J]. Remote Sensing of Environment, 1996, 58(3):257-266.
[10]
徐涵秋. 利用改进的归一化差异水体指数(MNDWI)提取水体信息的研究[J]. 遥感学报, 2005, 9(5): 589-595.
(Xu Hanqiu. A study on information extraction of water body with the modified normalized difference water index (MNDWI)[J]. Journal of Remote Sensing, 2005, 9(5): 589-595. (in Chinese))
[11]
Feyisa G L, Meilby H, Fensholt R, et al. Automated Water Extraction Index: A new technique for surface water mapping using Landsat imagery[J]. Remote Sensing of Environment, 2014, 140: 23-35.
[12]
洪亮, 黄雅君, 杨昆, 等. 复杂环境下高分二号遥感影像的城市地表水体提取[J]. 遥感学报, 2019, 23(5): 871-882.
(Hong Liang, Huang Yajun, Yang Kun, et al. Study on urban surface water extraction from heterogeneous environments using GF-2 remotely sensed images[J]. National Remote Sensing Bulletin, 2019, 23(5): 871-882. (in Chinese))
[13]
Guo Mingqiang, Zhang Haixue, Huang Ying, et al. Shadow removal method for high-resolution aerial remote sensing images based on region group matching[J]. Expert Systems with Applications, 2024, 255: 124739.
[14]
Liu Jiahang, Wang Xiaozhen, Mao Guo, et al. Shadow detection in remote sensing images based on spectral radiance separability enhancement[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(5): 3438-3449.
[15]
Li Yansheng, Dang Bo, Zhang Yongjun, et al. Water body classification from high-resolution optical remote sensing imagery: achievements and perspectives[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 187: 306-327.
[16]
梁敏, 汪西莉. 结合超分辨率和域适应的遥感图像语义分割方法[J]. 计算机学报, 2022, 45(12):2619-2636.
(Liang Min, Wang Xili. Semantic segmentation model for remote sensing images combining super resolution and domain adaption[J]. Chinese Journal of Computers, 2022, 45(12): 2619-2636. (in Chinese))
[17]
张云康, 刘懿, 肖婉, 等. 基于Segformer模型的洪水淹没范围提取与对比[J]. 长江科学院院报, 2025, 42(10): 165-173, 182.
Abstract
洪水期间,复杂地形和水域边界的快速变化极大增加了准确提取洪水淹没范围的难度。传统的人工和卫星遥感方法常因云层覆盖、光照变化和传感器限制等因素,导致其精度不足,且效率低下、成本高昂。基于此,利用深度学习与机器视觉技术,构建了针对洪水监测的“RiverDataset”数据集,并基于此数据集评估Segformer模型在洪水淹没范围提取中的效果。同时对比分析Segformer模型与基于ResNet50和VGG16的U-Net模型的水体分割性能。研究结果表明:Segformer模型凭借其Transformer结构,有效捕捉了广泛的上下文信息,并通过多尺度特征融合技术,保证了在复杂环境中的分割性能和信息的完整性;在平均交并比(mIoU)、平均像素准确率(mPA)、精确率和召回率等关键性能指标上,Segformer模型显著优于基于ResNet50和VGG16的U-Net模型。综上可知,Segformer模型在处理精确提取边界和有效排除干扰物的复杂分割任务中具有明显优势。
(Zhang Yunkang, Liu Yi, Xiao Wan, et al. Flood inundation range extraction and comparative analysis based on segformer model[J]. Journal of Changjiang River Scientific Research Institute, 2025, 42(10): 165-173, 182. (in Chinese))
[18]
王国杰, 胡一凡, 张森, 等. 深度卷积神经网络的遥感影像水体识别[J]. 遥感学报, 2022, 26(11):2304-2316.
(Wang Guojie, Hu Yifan, Zhang Sen, et al. Water identification from the GF-1 satellite image based on the deep convolutional neural networks[J]. Journal of Remote Sensing, 2022, 26(11): 2304-2316. (in Chinese))
[19]
Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]// 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 7-12, 2015, Boston, MA, USA. New York: IEEE Press, 2015: 3431-3440.
[20]
Ronneberger O, Fischer P, Brox T. U-Net: Convolutional networks for biomedical image segmentation[C]// Medical Image Computing and Computer-Assisted Intervention-MICCAI 2015. Cham: Springer, 2015: 234-241.
[21]
Farooq B, Manocha A. Small water body extraction in remote sensing with enhanced CNN architecture[J]. Applied Soft Computing, 2025, 169: 112544.
[22]
黎东丰, 陈雨人, 余博. 基于多层次特征融合的路面裂缝检测方法[J]. 计算机工程, 2026, 52(1):154-165.
(Li Dongfeng, Chen Yuren, Yu Bo. Pavement crack detection method based on multi-level feature fusion[J]. Computer Engineering, 2026, 52(1):154-165. (in Chinese))

Current U-Net-based pavement crack detection methods do not fully consider the interaction between the features of each level of the encoder, causing incomplete detection results or missed detections because of information loss during the downsampling process. To address this issue, this study proposes a pavement crack detection method based on multi-level feature fusion. In the encoding stage, the features of cracks at different levels are extracted to form crack feature representations from shallow to deep layers. In the skip connection section, a cross-level fusion strategy based on an improved Channel Cross Transformer (CCT) is adopted to enhance the complementarity between features at each level and enrich the expression of crack features. In the decoding stage, the feature fusion module is used to optimize the decoder's utilization of encoder features, promote the transmission of crack features, and improve the perception ability of crack features. In a series of comparative and ablation experiments on two public datasets, DeepCrack and CRACK500, the proposed method outperforms six other methods, including DeepCrack and Swin-UNet. On DeepCrack, the proposed method increases the F1 value by 2.30 and 2.51 percentage points, respectively, compared to those of DeepCrack and Swin-UNet, while on CRACK500, it increases by 1.65 and 1.00 percentage points, respectively.

[23]
孙静, 孙福奇, 郝世杰, 等. 一种基于小波变换的两阶段低照度图像增强方法[J]. 计算机学报, 2025, 48(5):1188-1211.
(Sun Jing, Sun Fuqi, Hao Shijie, et al. Two-stage low-light image enhancement based on wavelet transform[J]. Chinese Journal of Computers, 2025, 48(5): 1188-1211. (in Chinese))
[24]
吴淞, 蓝鑫, 单靖杨, 等. 基于注意力机制和多尺度融合的U-Net改进算法[J]. 计算机应用, 2024, 44(增刊2): 24-28.
(Wu Song, Lan Xin, Shan Jingyang, et al. Improved U-Net algorithm based on attention mechanism and multi-scale fusion[J]. Journal of Computer Applications, 2024, 44(S2): 24-28. (in Chinese))
[25]
Zhou Zongwei, Rahman Siddiquee M M, Tajbakhsh N, et al. UNet++: A nested U-Net architecture for medical image segmentation[C]// Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support. Cham: Springer, 2018: 3-11.
[26]
Chen L C, Zhu Y, Papandreou G, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation[C]// Computer Vision - ECCV 2018. Cham: Springer, 2018: 833-851.
[27]
Sun Ke, Zhao Yang, Jiang Borui, et al. High-resolution representations for labeling pixels and regions. arXiv 2019. arXiv preprint arXiv :1904.04514. 2019; 5(7).
[28]
徐丹青, 吴一全. 光学遥感图像目标检测的深度学习算法研究进展[J]. 遥感学报, 2024, 28(12): 3045-3073.
(Xu Danqing, Wu Yiquan. Progress of research on deep learning algorithms for object detection in optical remote sensing images[J]. National Remote Sensing Bulletin, 2024, 28(12): 3045-3073. (in Chinese))
[29]
Li Juanjuan, Wang Chao, Xu Lu, et al. Multitemporal water extraction of Dongting Lake and Poyang Lake based on an automatic water extraction and dynamic monitoring framework[J]. Remote Sensing, 2021, 13(5): 865.
[30]
Fu Jun, Liu Jing, Tian Haijie, et al. Dual attention network for scene segmentation[C]// 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 15-20, 2019, Long Beach, CA, USA. New York: IEEE Press, 2019: 3141-3149.
[31]
李振轩, 黄敏儿, 高飞, 等. 基于U-Net、U-Net++和Attention-U-Net网络的遥感影像水体提取[J]. 测绘通报, 2024(8):26-30.
Abstract
目前,深度学习在高分辨率遥感影像水体提取方面的应用已成为遥感领域的研究热点。其中基于U-Net网络的算法在水体提取中表现出较好的性能,但鲜有研究对不同U-Net网络算法在水体提取任务中的性能差异进行深入比较。因此,本文选择U-Net、U-Net++和Attention-U-Net 3种卷积神经网络,基于GID数据集,进行试验与定量分析。结果表明:U-Net++的训练精度最高,其次为U-Net、Attention-U-Net,三者分别为0.912、0.907、0.899;U-Net++的边缘提取能力优于其他两种网络;在分割不同类型水体和区分遥感影像中与水体区域相似的非水体区域上,U-Net++的提取效果显著,U-Net和Attention-U-Net易出现漏提现象,效果欠佳。
(Li Zhenxuan, Huang Miner, Gao Fei, et al. Remote sensing image water body extraction based on U-Net, U-Net++and Attention-U-Net networks[J]. Bulletin of Surveying and Mapping, 2024(8):26-30. (in Chinese))
Currently, the application of deep learning in the extraction of water bodies from high-resolution remote sensing images has become a research hotspot in the remote sensing field. Among them, algorithms based on the U-Net network have demonstrated good performance in water body extraction. However, there is scarce research that provides in-depth and detailed comparisons of the performance differences of different U-Net network algorithms in water body extraction tasks. Therefore, this article selects three convolutional neural networks, named U-Net, U-Net++, and Attention-U-Net, and based on the GID dataset, draws conclusions through experiments and quantitative analysis. The results indicate that: U-Net++ achieves the highest training accuracy, followed by U-Net and Attention-U-Net, with accuracies of 0.912, 0.907, and 0.899 respectively. U-Net++ exhibits superior edge extraction capability compared to the other two networks. In segmenting different types of water bodies and distinguishing non-water areas similar to water bodies in remote sensing images, U-Net++ shows significantly better extraction results, while U-Net and Attention-U-Net are prone to omission errors and exhibit suboptimal performance.
[32]
姜文文, 夏英. 改进U-Net的多尺度特征融合遥感图像语义分割网络[J]. 计算机科学, 2025, 52(5): 212-219.
(Jiang Wenwen, Xia Ying. Improved U-Net multi-scale feature fusion semantic segmentation network for remote sensing images[J]. Computer Science, 2025, 52(5): 212-219. (in Chinese))
[33]
温世雄, 智敏. 视觉Transformer在细粒度图像分类中的应用综述[J]. 计算机工程与应用, 2025, 61(23):24-37.
Abstract
细粒度图像分类(fine-grained image classification,FGIC)旨在识别视觉上高度相似但存在细微差异的子类别。随着深度学习的快速发展,FGIC算法已由传统强监督学习逐步发展至弱监督学习。视觉Transformer(ViT)凭借其多头自注意力机制,无须依赖手工标注,同时克服了基于卷积神经网络(CNN)算法在感受野和全局建模能力上的局限性,成为该任务的主流方法之一。对FGIC的特点与难点进行概述,简要介绍ViT的基本架构及其优势。根据不同的特征融合策略将基于ViT的改进算法分成层次、多局部及多粒度三种特征融合方法,对每类方法的改进方式进行详细的图示说明,并对各类技术方法的机制进行详细阐述和总结分析。梳理了常用的公开数据集,并根据当前研究的局限性提出未来的研究方向,以进一步挖掘ViT在细粒度图像分类任务中的应用潜力。
(Wen Shixiong, Zhi Min. Survey of vision transformers for fine-grained image classification[J]. Computer Engineering and Applications, 2025, 61(23):24-37. (in Chinese))
Fine-grained image classification (FGIC) aims to identify subcategories that are visually highly similar yet exhibit subtle differences. With the rapid advancement of deep learning, FGIC algorithms have gradually evolved from traditional fully supervised learning to weakly supervised approaches. Vision Transformers (ViTs), leveraging multi-head self-attention mechanisms, eliminate the reliance on manual annotations and overcome the limitations of convolutional neural networks (CNNs) in terms of receptive field size and global modeling capacity, becoming one of the mainstream methods for this task. This paper first outlines the key characteristics and challenges of FGIC, and briefly introduces the architecture and advantages of ViT. Based on different feature fusion strategies, existing ViT-based improvements are categorized into hierarchical fusion, multi-local fusion, and multi-granularity fusion. The modifications of each category are illustrated in detail, and their underlying mechanisms are systematically analyzed and summarized. In addition, commonly used public datasets are reviewed, and future research directions are proposed based on current limitations, aiming to further explore the potential of ViT in FGIC tasks.
[34]
董亦凡, 孙文礼, 赵洋, 等. 遥感图像半监督语义分割方法研究综述[J]. 计算机工程与应用, 2025, 61(24): 86-102.
Abstract
遥感图像语义分割作为遥感技术领域的重要研究方向,旨在对遥感图像进行像素级的分割并精确分类至预定义的地物类别中。传统方法依赖于大量的像素级标注数据,但人工标注耗时且费力。为解决这一问题,半监督学习方法被引入,通过结合少量标注样本和大量无标注数据进行模型训练,显著降低了标注需求。通过分析和总结近年来关于遥感图像半监督语义分割的相关研究,系统梳理现有方法的分类体系,并深入剖析其优势与不足:从核心思想和技术策略出发,对现有的遥感图像半监督语义分割方法进行了体系归纳和类别划分,并分别讨论了其创新与局限性;介绍了遥感图像半监督语义分割研究广泛使用的数据集;基于常用实验设置和评价指标,在不同数据集上开展了多方法对比分析;探讨了遥感图像半监督语义分割的未来研究趋势。
(Dong Yifan, Sun Wenli, Zhao Yang, et al. Survey on semi-supervised semantic segmentation methods for remote sensing images[J]. Computer Engineering and Applications, 2025, 61(24): 86-102. (in Chinese))
As an important research direction in the field of remote sensing technology, semantic segmentation of remote sensing images aims to segment remote sensing images at the pixel level and accurately classify them into predefined ground object categories. Traditional methods rely on a large number of pixel-level annotation data, but manual annotation is time-consuming and laborious. In order to solve this problem, a semi-supervised learning method is introduced. By combining a small number of labeled samples and a large number of unlabeled data for model training, the labeling requirements are significantly reduced. By analyzing and summarizing the related research on semi-supervised semantic segmentation of remote sensing images in recent years, the classification system of existing methods is systematically sorted out, and its advantages and disadvantages are deeply analyzed. Starting from the core ideas and technical strategies, the existing semi-supervised semantic segmentation methods of remote sensing images are systematically summarized and classified, and their innovations and limitations are discussed respectively. The datasets widely used in the research of semi-supervised semantic segmentation of remote sensing images are introduced. Based on the commonly used experimental settings and evaluation indicators, multi-method comparative analysis is carried out on different datasets. Finally, the future research trends of semi-supervised semantic segmentation of remote sensing images are discussed.
[35]
Hou Qibin, Zhang Li, Cheng Mingming, et al. Strip pooling: Rethinking spatial pooling for scene parsing[C]// 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 13-19, 2020. Seattle, WA, USA. New York: IEEE Press, 2020: 4002-4011.
[36]
Yang Lingxiao, Zhang Ruyuan, Li Lida, et al. Simam: A simple, parameter-free attention module for convolutional neural networks[C]// International conference on machine learning. Proceedings of Machine Learning Research 139: 11863-11874.
[37]
Yin Yuheng, Guo Yanggang, Deng Liwei, et al. Improved PSPNet-based water shoreline detection in complex inland river scenarios[J]. Complex & Intelligent Systems, 2023, 9(1):233-245.
[38]
Ren Qiuyu, Lu Zhiying, Wu Haopeng, et al. HR-Net:A landmark based high realistic face reenactment network[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 33(11): 6347-6359.
[39]
Shi Sunan, Zhong Yanfei, Liu Yinhe, et al. Multi-temporal urban semantic understanding based on GF-2 remote sensing imagery: from tri-temporal datasets to multi-task mapping[J]. International Journal of Digital Earth, 2023, 16(1): 3321-3347.
PDF(3114 KB)

Accesses

Citation

Detail

Sections
Recommended

/