| ||||
| ||||
![]() Title:Asymmetric Hierarchical Fusion for Visual Grounding in SAR Images Conference:ACIIDS2026 Tags:Detection transformer, Hierarchical fusion, Multimodal learning, Remote sensing visual grounding and Synthetic aperture radar (SAR) Abstract: Visual grounding on synthetic aperture radar (SAR) imagery aims to localize objects described by natural language expressions. However, existing SAR grounding studies remain limited to single-category scenarios and lack relational descriptions, preventing models from generalizing to realistic multi-object scenes. To address this issue, we introduce SAAVG, a new multi-class SAR visual grounding dataset built from multiple high-resolution SAR detection sources and enriched with automatically generated and manually verified referring expressions. Building upon the LQVG framework, we propose an Asymmetric Hierarchical Fusion mechanism that deepens the Vision–Language Interaction pathway through iterative cross-modal refinement, motivated by the dominant role of this pathway in visual grounding performance. We examine two variants, Shared-Weight and Per-Scale, to characterize their behavior across SAR and optical domains. Experiments on DIOR-RSVG, SARVG1.0, and SAAVG show that the Shared-Weight variant achieves state-of-the-art results on SARVG1.0 (Pr@0.5 92.01%, mIoU 83.39%), while the Per-Scale design performs better in optical imagery, revealing modality-dependent fusion dynamics. The proposed dataset and findings establish a useful benchmark and offer valuable insights into hierarchical cross-modal interaction for SAR visual grounding. Asymmetric Hierarchical Fusion for Visual Grounding in SAR Images ![]() Asymmetric Hierarchical Fusion for Visual Grounding in SAR Images | ||||
| Copyright © 2002 – 2026 EasyChair |
