physfusion: A Transformer-based Dual-Stream Radar and Vision Fusion Framework for Open Water Surface Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918366626709504 |
|---|---|
| author | Wan, Yuting Sun, Liguo Hao, Jiuwu Zhang, Zao LV, Pin |
| author_facet | Wan, Yuting Sun, Liguo Hao, Jiuwu Zhang, Zao LV, Pin |
| contents | Detecting water-surface targets for Unmanned Surface Vehicles (USVs) is challenging due to wave clutter,
specular reflections, and weak appearance cues in long-range observations. Although 4D millimeter-wave
radar complements cameras under degraded illumination, maritime radar point clouds are sparse and
intermittent, with reflectivity attributes exhibiting heavy-tailed variations under scattering and
multipath, making conventional fusion designs struggle to exploit radar cues effectively.
We propose PhysFusion, a physics-informed radar-image detection framework for water-surface perception.
The framework integrates: (1) a Physics-Informed Radar Encoder (PIR Encoder) with an RCS Mapper and
Quality Gate, transforming per-point radar attributes into compact scattering priors and predicting
point-wise reliability for robust feature learning under clutter; (2) a Radar-guided Interactive Fusion
Module (RIFM) performing query-level radar-image fusion between semantically enriched radar features and
multi-scale visual features, with the radar branch modeled by a dual-stream backbone including a
point-based local stream and a transformer-based global stream using Scattering-Aware Self-Attention
(SASA); and (3) a Temporal Query Aggregation module (TQA) aggregating frame-wise fused queries over a
short temporal window for temporally consistent representations.
Experiments on WaterScenes and FLOW demonstrate that PhysFusion achieves 59.7% mAP50:95 and 90.3% mAP50
on WaterScenes (T=5 radar history) using 5.6M parameters and 12.5G FLOPs, and reaches 94.8% mAP50 and
46.2% mAP50:95 on FLOW under radar+camera setting. Ablation studies quantify the contributions of PIR
Encoder, SASA-based global reasoning, and RIFM. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_01947 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | physfusion: A Transformer-based Dual-Stream Radar and Vision Fusion Framework for Open Water Surface Object Detection Wan, Yuting Sun, Liguo Hao, Jiuwu Zhang, Zao LV, Pin Computer Vision and Pattern Recognition Artificial Intelligence Detecting water-surface targets for Unmanned Surface Vehicles (USVs) is challenging due to wave clutter, specular reflections, and weak appearance cues in long-range observations. Although 4D millimeter-wave radar complements cameras under degraded illumination, maritime radar point clouds are sparse and intermittent, with reflectivity attributes exhibiting heavy-tailed variations under scattering and multipath, making conventional fusion designs struggle to exploit radar cues effectively. We propose PhysFusion, a physics-informed radar-image detection framework for water-surface perception. The framework integrates: (1) a Physics-Informed Radar Encoder (PIR Encoder) with an RCS Mapper and Quality Gate, transforming per-point radar attributes into compact scattering priors and predicting point-wise reliability for robust feature learning under clutter; (2) a Radar-guided Interactive Fusion Module (RIFM) performing query-level radar-image fusion between semantically enriched radar features and multi-scale visual features, with the radar branch modeled by a dual-stream backbone including a point-based local stream and a transformer-based global stream using Scattering-Aware Self-Attention (SASA); and (3) a Temporal Query Aggregation module (TQA) aggregating frame-wise fused queries over a short temporal window for temporally consistent representations. Experiments on WaterScenes and FLOW demonstrate that PhysFusion achieves 59.7% mAP50:95 and 90.3% mAP50 on WaterScenes (T=5 radar history) using 5.6M parameters and 12.5G FLOPs, and reaches 94.8% mAP50 and 46.2% mAP50:95 on FLOW under radar+camera setting. Ablation studies quantify the contributions of PIR Encoder, SASA-based global reasoning, and RIFM. |
| title | physfusion: A Transformer-based Dual-Stream Radar and Vision Fusion Framework for Open Water Surface Object Detection |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2603.01947 |