MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Minghui, Huang, Wei-Hsing, Liang, Shaofeng, Liu, Daizong, Wen, Tai-Hao, Wang, Gang, Guan, Runwei, Ding, Weiping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
by: Cao, Zhixiang, et al.
Published: (2026)
by: Cao, Zhixiang, et al.
Published: (2026)
OptiPMB: Enhancing 3D Multi-Object Tracking with Optimized Poisson Multi-Bernoulli Filtering
by: Ding, Guanhua, et al.
Published: (2025)
by: Ding, Guanhua, et al.
Published: (2025)
WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
by: Yao, Shanliang, et al.
Published: (2023)
by: Yao, Shanliang, et al.
Published: (2023)
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
by: Chen, Chenglizhao, et al.
Published: (2025)
by: Chen, Chenglizhao, et al.
Published: (2025)
Cognitive Disentanglement for Referring Multi-Object Tracking
by: Liang, Shaofeng, et al.
Published: (2025)
by: Liang, Shaofeng, et al.
Published: (2025)
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
by: Guan, Runwei, et al.
Published: (2026)
by: Guan, Runwei, et al.
Published: (2026)
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
Learning When to See and When to Feel: Adaptive Vision-Torque Fusion for Contact-Aware Manipulation
by: Lei, Jiuzhou, et al.
Published: (2026)
by: Lei, Jiuzhou, et al.
Published: (2026)
Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
by: Guan, Runwei, et al.
Published: (2023)
by: Guan, Runwei, et al.
Published: (2023)
RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
by: Zhang, Zhang, et al.
Published: (2025)
by: Zhang, Zhang, et al.
Published: (2025)
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
by: Sun, Xiaolou, et al.
Published: (2026)
by: Sun, Xiaolou, et al.
Published: (2026)
4D-CAAL: 4D Radar-Camera Calibration and Auto-Labeling for Autonomous Driving
by: Yao, Shanliang, et al.
Published: (2026)
by: Yao, Shanliang, et al.
Published: (2026)
MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding
by: Liang, Jing, et al.
Published: (2025)
by: Liang, Jing, et al.
Published: (2025)
Beyond Uncertainty: Risk-Aware Active View Acquisition for Safe Robot Navigation and 3D Scene Understanding with FisherRF
by: Liu, Guangyi, et al.
Published: (2024)
by: Liu, Guangyi, et al.
Published: (2024)
SpotLight: Robotic Scene Understanding through Interaction and Affordance Detection
by: Engelbracht, Tim, et al.
Published: (2024)
by: Engelbracht, Tim, et al.
Published: (2024)
PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
by: Soleymanzadeh, Davood, et al.
Published: (2025)
by: Soleymanzadeh, Davood, et al.
Published: (2025)
Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene Understanding
by: Tang, Yifan, et al.
Published: (2024)
by: Tang, Yifan, et al.
Published: (2024)
OmniVIC: A Self-Improving Variable Impedance Controller with Vision-Language In-Context Learning for Safe Robotic Manipulation
by: Zhang, Heng, et al.
Published: (2025)
by: Zhang, Heng, et al.
Published: (2025)
ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent Manipulation
by: Wang, Enyi, et al.
Published: (2026)
by: Wang, Enyi, et al.
Published: (2026)
Active Vision for Scene Understanding
by: Grotz, Markus
Published: (2022)
by: Grotz, Markus
Published: (2022)
SELF-VLA: A Skill Enhanced Agentic Vision-Language-Action Framework for Contact-Rich Disassembly
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Fusion Dynamical Systems with Machine Learning in Imitation Learning: A Comprehensive Overview
by: Hu, Yingbai, et al.
Published: (2024)
by: Hu, Yingbai, et al.
Published: (2024)
AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding
by: Lin, Fei, et al.
Published: (2025)
by: Lin, Fei, et al.
Published: (2025)
Exosense: A Vision-Based Scene Understanding System For Exoskeletons
by: Wang, Jianeng, et al.
Published: (2024)
by: Wang, Jianeng, et al.
Published: (2024)
POIROT: Investigating Direct Tangible vs. Digitally Mediated Interaction and Attitude Moderation in Multi-party Murder Mystery Games
by: Chen, Wen, et al.
Published: (2026)
by: Chen, Wen, et al.
Published: (2026)
MorphoCopter: Design, Modeling, and Control of a New Transformable Quad-Bi Copter
by: Modi, Harsh, et al.
Published: (2025)
by: Modi, Harsh, et al.
Published: (2025)
HELIOS: Hierarchical Exploration for Language-Grounded Interaction in Open Scenes
by: Ashton, Katrina, et al.
Published: (2025)
by: Ashton, Katrina, et al.
Published: (2025)
USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
by: Yao, Shanliang, et al.
Published: (2025)
by: Yao, Shanliang, et al.
Published: (2025)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Interleaved LLM and Motion Planning for Generalized Multi-Object Collection in Large Scene Graphs
by: Yang, Ruochu, et al.
Published: (2025)
by: Yang, Ruochu, et al.
Published: (2025)
OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding
by: Zhu, Siting, et al.
Published: (2026)
by: Zhu, Siting, et al.
Published: (2026)
Scene-Adaptive Motion Planning with Explicit Mixture of Experts and Interaction-Oriented Optimization
by: Zhu, Hongbiao, et al.
Published: (2025)
by: Zhu, Hongbiao, et al.
Published: (2025)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
Scene Exploration by Vision-Language Models
by: Sripada, Venkatesh, et al.
Published: (2024)
by: Sripada, Venkatesh, et al.
Published: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
by: Jia, Baoxiong, et al.
Published: (2024)
by: Jia, Baoxiong, et al.
Published: (2024)
DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving
by: Hou, Xinmeng, et al.
Published: (2025)
by: Hou, Xinmeng, et al.
Published: (2025)
GRID: Scene-Graph-based Instruction-driven Robotic Task Planning
by: Ni, Zhe, et al.
Published: (2023)
by: Ni, Zhe, et al.
Published: (2023)
Robotic Scene Cloning:Advancing Zero-Shot Robotic Scene Adaptation in Manipulation via Visual Prompt Editing
by: Huang, Binyuan, et al.
Published: (2026)
by: Huang, Binyuan, et al.
Published: (2026)
Similar Items
-
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
by: Cao, Zhixiang, et al.
Published: (2026) -
OptiPMB: Enhancing 3D Multi-Object Tracking with Optimized Poisson Multi-Bernoulli Filtering
by: Ding, Guanhua, et al.
Published: (2025) -
WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
by: Yao, Shanliang, et al.
Published: (2023) -
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
by: Guan, Runwei, et al.
Published: (2025) -
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
by: Chen, Chenglizhao, et al.
Published: (2025)