RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Zhijian, Feng, Chengjian, Yan, Feng, Xiao, Baihui, Jie, Zequn, Zhong, Yujie, Liang, Xiaodan, Ma, Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
di: Yan, Feng, et al.
Pubblicazione: (2024)
di: Yan, Feng, et al.
Pubblicazione: (2024)
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
di: Xiao, Baihui, et al.
Pubblicazione: (2025)
di: Xiao, Baihui, et al.
Pubblicazione: (2025)
RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction
di: Zhong, Yufeng, et al.
Pubblicazione: (2025)
di: Zhong, Yufeng, et al.
Pubblicazione: (2025)
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
di: Guo, Ziang, et al.
Pubblicazione: (2026)
di: Guo, Ziang, et al.
Pubblicazione: (2026)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
di: Min, Chen, et al.
Pubblicazione: (2023)
di: Min, Chen, et al.
Pubblicazione: (2023)
NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving
di: Peng, Qucheng, et al.
Pubblicazione: (2025)
di: Peng, Qucheng, et al.
Pubblicazione: (2025)
U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks
di: Feng, Tongtong, et al.
Pubblicazione: (2025)
di: Feng, Tongtong, et al.
Pubblicazione: (2025)
Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency
di: Zarghani, Abolfazl, et al.
Pubblicazione: (2025)
di: Zarghani, Abolfazl, et al.
Pubblicazione: (2025)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
di: Zheng, Liming, et al.
Pubblicazione: (2024)
di: Zheng, Liming, et al.
Pubblicazione: (2024)
WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
di: Liu, Yanbaihui, et al.
Pubblicazione: (2024)
di: Liu, Yanbaihui, et al.
Pubblicazione: (2024)
BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
di: Tan, Wentao, et al.
Pubblicazione: (2025)
di: Tan, Wentao, et al.
Pubblicazione: (2025)
CineWild: Balancing Art and Robotics for Ethical Wildlife Documentary Filmmaking
di: Pueyo, Pablo, et al.
Pubblicazione: (2025)
di: Pueyo, Pablo, et al.
Pubblicazione: (2025)
Bringing Robots Home: The Rise of AI Robots in Consumer Electronics
di: Dong, Haiwei, et al.
Pubblicazione: (2024)
di: Dong, Haiwei, et al.
Pubblicazione: (2024)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
di: Yu, Wenda, et al.
Pubblicazione: (2026)
di: Yu, Wenda, et al.
Pubblicazione: (2026)
A Multimedia Framework for Continuum Robots: Systematic, Computational, and Control Perspectives
di: Hsieh, Po-Yu, et al.
Pubblicazione: (2024)
di: Hsieh, Po-Yu, et al.
Pubblicazione: (2024)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
di: Xu, Siyuan, et al.
Pubblicazione: (2026)
di: Xu, Siyuan, et al.
Pubblicazione: (2026)
RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System
di: Choudhury, Nitin, et al.
Pubblicazione: (2026)
di: Choudhury, Nitin, et al.
Pubblicazione: (2026)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
di: Huang, Yiheng, et al.
Pubblicazione: (2025)
di: Huang, Yiheng, et al.
Pubblicazione: (2025)
CALMM-Drive: Confidence-Aware Autonomous Driving with Large Multimodal Model
di: Yao, Ruoyu, et al.
Pubblicazione: (2024)
di: Yao, Ruoyu, et al.
Pubblicazione: (2024)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
di: Liu, Fanfan, et al.
Pubblicazione: (2024)
di: Liu, Fanfan, et al.
Pubblicazione: (2024)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
di: Wang, Sen, et al.
Pubblicazione: (2025)
di: Wang, Sen, et al.
Pubblicazione: (2025)
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
di: Qian, Kangan, et al.
Pubblicazione: (2026)
di: Qian, Kangan, et al.
Pubblicazione: (2026)
Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy
di: Mo, Feng, et al.
Pubblicazione: (2024)
di: Mo, Feng, et al.
Pubblicazione: (2024)
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving
di: Hou, Xinmeng, et al.
Pubblicazione: (2025)
di: Hou, Xinmeng, et al.
Pubblicazione: (2025)
MotiBo: The Impact of Interactive Digital Storytelling Robots on Student Motivation through Self-Determination Theory
di: Fung, Ka Yan, et al.
Pubblicazione: (2026)
di: Fung, Ka Yan, et al.
Pubblicazione: (2026)
Flight Patterns for Swarms of Drones
di: Zhu, Shuqin, et al.
Pubblicazione: (2024)
di: Zhu, Shuqin, et al.
Pubblicazione: (2024)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
di: Sun, Qiao, et al.
Pubblicazione: (2025)
di: Sun, Qiao, et al.
Pubblicazione: (2025)
LMMCoDrive: Cooperative Driving with Large Multimodal Model
di: Liu, Haichao, et al.
Pubblicazione: (2024)
di: Liu, Haichao, et al.
Pubblicazione: (2024)
Scaling Spatial Intelligence with Multimodal Foundation Models
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
di: Yang, Xiaocui, et al.
Pubblicazione: (2024)
di: Yang, Xiaocui, et al.
Pubblicazione: (2024)
The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition
di: Kong, Lingdong, et al.
Pubblicazione: (2024)
di: Kong, Lingdong, et al.
Pubblicazione: (2024)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
di: An, Xiao, et al.
Pubblicazione: (2026)
di: An, Xiao, et al.
Pubblicazione: (2026)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
di: Sun, Hao, et al.
Pubblicazione: (2024)
di: Sun, Hao, et al.
Pubblicazione: (2024)
Evaluating Magic Leap 2 Tool Tracking for AR Sensor Guidance in Industrial Inspections
di: Masuhr, Christian, et al.
Pubblicazione: (2025)
di: Masuhr, Christian, et al.
Pubblicazione: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
RoboCar: A Rapidly Deployable Open-Source Platform for Autonomous Driving Research
di: Testouri, Mehdi, et al.
Pubblicazione: (2024)
di: Testouri, Mehdi, et al.
Pubblicazione: (2024)
WoW: Towards a World omniscient World model Through Embodied Interaction
di: Chi, Xiaowei, et al.
Pubblicazione: (2025)
di: Chi, Xiaowei, et al.
Pubblicazione: (2025)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
di: Zhong, Zesen, et al.
Pubblicazione: (2025)
di: Zhong, Zesen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
di: Yan, Feng, et al.
Pubblicazione: (2024) -
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
di: Xiao, Baihui, et al.
Pubblicazione: (2025) -
RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction
di: Zhong, Yufeng, et al.
Pubblicazione: (2025) -
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
di: Guo, Ziang, et al.
Pubblicazione: (2026) -
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
di: Min, Chen, et al.
Pubblicazione: (2023)