WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914353098260480 |
|---|---|
| author | Guan, Runwei Liang, Shaofeng Ouyang, Ningwei Fei, Weichen Yao, Shanliang Dai, Wei Ge, Chenhao Sun, Penglei Zhu, Xiaohui Huang, Tao Liu, Ryan Wen Xiong, Hui |
| author_facet | Guan, Runwei Liang, Shaofeng Ouyang, Ningwei Fei, Weichen Yao, Shanliang Dai, Wei Ge, Chenhao Sun, Penglei Zhu, Xiaohui Huang, Tao Liu, Ryan Wen Xiong, Hui |
| contents | While autonomous navigation has achieved remarkable success in passive perception (e.g., object detection and segmentation), it remains fundamentally constrained by a void in knowledge-driven, interactive environmental cognition. In the high-stakes domain of maritime navigation, the ability to bridge the gap between raw visual perception and complex cognitive reasoning is not merely an enhancement but a critical prerequisite for Autonomous Surface Vessels to execute safe and precise maneuvers. To this end, we present WaterVideoQA, the first large-scale, comprehensive Video Question Answering benchmark specifically engineered for all-waterway environments. This benchmark encompasses 3,029 video clips across six distinct waterway categories, integrating multifaceted variables such as volatile lighting and dynamic weather to rigorously stress-test ASV capabilities across a five-tier hierarchical cognitive framework. Furthermore, we introduce NaviMind, a pioneering multi-agent neuro-symbolic system designed for open-ended maritime reasoning. By synergizing Adaptive Semantic Routing, Situation-Aware Hierarchical Reasoning, and Autonomous Self-Reflective Verification, NaviMind transitions ASVs from superficial pattern matching to regulation-compliant, interpretable decision-making. Experimental results demonstrate that our framework significantly transcends existing baselines, establishing a new paradigm for intelligent, trustworthy interaction in dynamic maritime environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_22923 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents Guan, Runwei Liang, Shaofeng Ouyang, Ningwei Fei, Weichen Yao, Shanliang Dai, Wei Ge, Chenhao Sun, Penglei Zhu, Xiaohui Huang, Tao Liu, Ryan Wen Xiong, Hui Computer Vision and Pattern Recognition Robotics While autonomous navigation has achieved remarkable success in passive perception (e.g., object detection and segmentation), it remains fundamentally constrained by a void in knowledge-driven, interactive environmental cognition. In the high-stakes domain of maritime navigation, the ability to bridge the gap between raw visual perception and complex cognitive reasoning is not merely an enhancement but a critical prerequisite for Autonomous Surface Vessels to execute safe and precise maneuvers. To this end, we present WaterVideoQA, the first large-scale, comprehensive Video Question Answering benchmark specifically engineered for all-waterway environments. This benchmark encompasses 3,029 video clips across six distinct waterway categories, integrating multifaceted variables such as volatile lighting and dynamic weather to rigorously stress-test ASV capabilities across a five-tier hierarchical cognitive framework. Furthermore, we introduce NaviMind, a pioneering multi-agent neuro-symbolic system designed for open-ended maritime reasoning. By synergizing Adaptive Semantic Routing, Situation-Aware Hierarchical Reasoning, and Autonomous Self-Reflective Verification, NaviMind transitions ASVs from superficial pattern matching to regulation-compliant, interpretable decision-making. Experimental results demonstrate that our framework significantly transcends existing baselines, establishing a new paradigm for intelligent, trustworthy interaction in dynamic maritime environments. |
| title | WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2602.22923 |