WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guan, Runwei, Liang, Shaofeng, Ouyang, Ningwei, Fei, Weichen, Yao, Shanliang, Dai, Wei, Ge, Chenhao, Sun, Penglei, Zhu, Xiaohui, Huang, Tao, Liu, Ryan Wen, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
von: Yao, Shanliang, et al.
Veröffentlicht: (2025)
von: Yao, Shanliang, et al.
Veröffentlicht: (2025)
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2023)
von: Guan, Runwei, et al.
Veröffentlicht: (2023)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
von: Wen, Jiawen, et al.
Veröffentlicht: (2026)
von: Wen, Jiawen, et al.
Veröffentlicht: (2026)
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving
von: Yao, Shanliang, et al.
Veröffentlicht: (2026)
von: Yao, Shanliang, et al.
Veröffentlicht: (2026)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026)
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026)
Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
RC-GeoCP: Geometric Consensus for Radar-Camera Collaborative Perception
von: Bai, Xiaokai, et al.
Veröffentlicht: (2026)
von: Bai, Xiaokai, et al.
Veröffentlicht: (2026)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
Cognitive Disentanglement for Referring Multi-Object Tracking
von: Liang, Shaofeng, et al.
Veröffentlicht: (2025)
von: Liang, Shaofeng, et al.
Veröffentlicht: (2025)
Human-Centric Goal Reasoning with Ripple-Down Rules
von: Brameld, Kenji, et al.
Veröffentlicht: (2024)
von: Brameld, Kenji, et al.
Veröffentlicht: (2024)
Video Spatial Reasoning with Object-Centric 3D Rollout
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Explore Human Parsing Modality for Action Recognition
von: Liu, Jinfu, et al.
Veröffentlicht: (2024)
von: Liu, Jinfu, et al.
Veröffentlicht: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
von: Vishal, Joseph Raj, et al.
Veröffentlicht: (2025)
von: Vishal, Joseph Raj, et al.
Veröffentlicht: (2025)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
Exploring Radar Data Representations in Autonomous Driving: A Comprehensive Review
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
von: Yao, Shanliang, et al.
Veröffentlicht: (2023)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
Reasoning-Enhanced Object-Centric Learning for Videos
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
von: Kovalev, Vsevolod, et al.
Veröffentlicht: (2025)
von: Kovalev, Vsevolod, et al.
Veröffentlicht: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
The Multi-Round Diagnostic RAG Framework for Emulating Clinical Reasoning
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
Representation and Characterization of Quasistationary Distributions for Markov Chains
von: Ben-Ari, Iddo, et al.
Veröffentlicht: (2024)
von: Ben-Ari, Iddo, et al.
Veröffentlicht: (2024)
RISE-Video: Can Video Generators Decode Implicit World Rules?
von: Liu, Mingxin, et al.
Veröffentlicht: (2026)
von: Liu, Mingxin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
von: Guan, Runwei, et al.
Veröffentlicht: (2025) -
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
von: Yao, Shanliang, et al.
Veröffentlicht: (2025) -
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
von: Guan, Runwei, et al.
Veröffentlicht: (2025) -
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2023) -
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)