Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Accurate semantic segmentation of high-resolution remote sensing images remains challenging due to complex scene structures, modality misalignment, and indistinct object boundaries. To address these challenges, we propose BMTransUNet, a boundary-aware gated multimodal Transformer U-Net, which introduces a unified fusion framework throughout the entire network to achieve effective RGB–DSM integration. Specifically, a dual-branch CNN encoder extracts multi-scale modality representations that are adaptively integrated via a Modality-Aware Gated Fusion (MAGF) module employing channel attention and semantic-guided gating. The fused features are subsequently refined through a Vision Transformer to capture long-range contextual dependencies. In the decoding stage, a Dual-branch ASPP (Dual-ASPP) module aggregates modality-specific contextual information, while an edge enhancement module reinforces boundary perception. Furthermore, an Edge-Aware Skip Connection (EASC) injects high-resolution structural cues to recover fine-grained details. Finally, a dual-head supervision strategy jointly optimizes semantic and edge predictions, thereby enforcing boundary consistency. Extensive experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate that BMTransUNet achieves superior segmentation accuracy and strong generalization capability compared with state-of-the-art methods.</p>

Show More

Keywords

module semantic segmentation highresolution modality

Related Articles

PORE

About

Connect