WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, Runwei, Liang, Shaofeng, Ouyang, Ningwei, Fei, Weichen, Yao, Shanliang, Dai, Wei, Ge, Chenhao, Sun, Penglei, Zhu, Xiaohui, Huang, Tao, Liu, Ryan Wen, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
by: Yao, Shanliang, et al.
Published: (2025)
by: Yao, Shanliang, et al.
Published: (2025)
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
by: Guan, Runwei, et al.
Published: (2023)
by: Guan, Runwei, et al.
Published: (2023)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
by: Wen, Jiawen, et al.
Published: (2026)
by: Wen, Jiawen, et al.
Published: (2026)
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving
by: Yao, Shanliang, et al.
Published: (2026)
by: Yao, Shanliang, et al.
Published: (2026)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
by: Li, Shaoxuan, et al.
Published: (2026)
by: Li, Shaoxuan, et al.
Published: (2026)
Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
by: Ma, Jianzhe, et al.
Published: (2026)
by: Ma, Jianzhe, et al.
Published: (2026)
RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception
by: Bai, Xiaokai, et al.
Published: (2026)
by: Bai, Xiaokai, et al.
Published: (2026)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
by: Wang, Chendong, et al.
Published: (2025)
by: Wang, Chendong, et al.
Published: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
Cognitive Disentanglement for Referring Multi-Object Tracking
by: Liang, Shaofeng, et al.
Published: (2025)
by: Liang, Shaofeng, et al.
Published: (2025)
Human-Centric Goal Reasoning with Ripple-Down Rules
by: Brameld, Kenji, et al.
Published: (2024)
by: Brameld, Kenji, et al.
Published: (2024)
Video Spatial Reasoning with Object-Centric 3D Rollout
by: Tang, Haoran, et al.
Published: (2025)
by: Tang, Haoran, et al.
Published: (2025)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Explore Human Parsing Modality for Action Recognition
by: Liu, Jinfu, et al.
Published: (2024)
by: Liu, Jinfu, et al.
Published: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
by: Wang, Xucheng, et al.
Published: (2026)
by: Wang, Xucheng, et al.
Published: (2026)
WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
by: Yao, Shanliang, et al.
Published: (2023)
by: Yao, Shanliang, et al.
Published: (2023)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
by: Vishal, Joseph Raj, et al.
Published: (2025)
by: Vishal, Joseph Raj, et al.
Published: (2025)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Exploring Radar Data Representations in Autonomous Driving: A Comprehensive Review
by: Yao, Shanliang, et al.
Published: (2023)
by: Yao, Shanliang, et al.
Published: (2023)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026)
by: Zhou, Zhiyu, et al.
Published: (2026)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
by: Hou, Minghui, et al.
Published: (2025)
by: Hou, Minghui, et al.
Published: (2025)
CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
by: Kovalev, Vsevolod, et al.
Published: (2025)
by: Kovalev, Vsevolod, et al.
Published: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
The Multi-Round Diagnostic RAG Framework for Emulating Clinical Reasoning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
Representation and Characterization of Quasistationary Distributions for Markov Chains
by: Ben-Ari, Iddo, et al.
Published: (2024)
by: Ben-Ari, Iddo, et al.
Published: (2024)
RISE-Video: Can Video Generators Decode Implicit World Rules?
by: Liu, Mingxin, et al.
Published: (2026)
by: Liu, Mingxin, et al.
Published: (2026)
Similar Items
-
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
by: Guan, Runwei, et al.
Published: (2025) -
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
by: Yao, Shanliang, et al.
Published: (2025) -
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
by: Guan, Runwei, et al.
Published: (2025) -
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
by: Guan, Runwei, et al.
Published: (2023) -
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)