RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Feng, Liu, Fanfan, Zheng, Liming, Zhong, Yufeng, Huang, Yiyang, Guan, Zechao, Feng, Chengjian, Ma, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
by: Xiao, Baihui, et al.
Published: (2025)
by: Xiao, Baihui, et al.
Published: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
by: Zheng, Liming, et al.
Published: (2024)
by: Zheng, Liming, et al.
Published: (2024)
Boosting Robotic Manipulation Generalization with Minimal Costly Data
by: Zheng, Liming, et al.
Published: (2025)
by: Zheng, Liming, et al.
Published: (2025)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026)
by: Yu, Wenda, et al.
Published: (2026)
Bringing Robots Home: The Rise of AI Robots in Consumer Electronics
by: Dong, Haiwei, et al.
Published: (2024)
by: Dong, Haiwei, et al.
Published: (2024)
CineWild: Balancing Art and Robotics for Ethical Wildlife Documentary Filmmaking
by: Pueyo, Pablo, et al.
Published: (2025)
by: Pueyo, Pablo, et al.
Published: (2025)
A Multimedia Framework for Continuum Robots: Systematic, Computational, and Control Perspectives
by: Hsieh, Po-Yu, et al.
Published: (2024)
by: Hsieh, Po-Yu, et al.
Published: (2024)
WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
by: Liu, Yanbaihui, et al.
Published: (2024)
by: Liu, Yanbaihui, et al.
Published: (2024)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
by: Sun, Qiao, et al.
Published: (2025)
by: Sun, Qiao, et al.
Published: (2025)
MotiBo: The Impact of Interactive Digital Storytelling Robots on Student Motivation through Self-Determination Theory
by: Fung, Ka Yan, et al.
Published: (2026)
by: Fung, Ka Yan, et al.
Published: (2026)
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
by: Guo, Ziang, et al.
Published: (2026)
by: Guo, Ziang, et al.
Published: (2026)
BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
by: Tan, Wentao, et al.
Published: (2025)
by: Tan, Wentao, et al.
Published: (2025)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
by: Xu, Siyuan, et al.
Published: (2026)
by: Xu, Siyuan, et al.
Published: (2026)
RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System
by: Choudhury, Nitin, et al.
Published: (2026)
by: Choudhury, Nitin, et al.
Published: (2026)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks
by: Feng, Tongtong, et al.
Published: (2025)
by: Feng, Tongtong, et al.
Published: (2025)
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
by: Qian, Kangan, et al.
Published: (2026)
by: Qian, Kangan, et al.
Published: (2026)
SemEval-2024 Task 3: Multimodal Emotion Cause Analysis in Conversations
by: Wang, Fanfan, et al.
Published: (2024)
by: Wang, Fanfan, et al.
Published: (2024)
Flight Patterns for Swarms of Drones
by: Zhu, Shuqin, et al.
Published: (2024)
by: Zhu, Shuqin, et al.
Published: (2024)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
by: Shirian, Melika, et al.
Published: (2025)
by: Shirian, Melika, et al.
Published: (2025)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
by: Yang, Xiaocui, et al.
Published: (2024)
by: Yang, Xiaocui, et al.
Published: (2024)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Evaluating Magic Leap 2 Tool Tracking for AR Sensor Guidance in Industrial Inspections
by: Masuhr, Christian, et al.
Published: (2025)
by: Masuhr, Christian, et al.
Published: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
WoW: Towards a World omniscient World model Through Embodied Interaction
by: Chi, Xiaowei, et al.
Published: (2025)
by: Chi, Xiaowei, et al.
Published: (2025)
Scaling Spatial Intelligence with Multimodal Foundation Models
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
by: Zhong, Zesen, et al.
Published: (2025)
by: Zhong, Zesen, et al.
Published: (2025)
One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
by: Wang, Sicheng, et al.
Published: (2025)
by: Wang, Sicheng, et al.
Published: (2025)
Personalized Image Generation with Large Multimodal Models
by: Xu, Yiyan, et al.
Published: (2024)
by: Xu, Yiyan, et al.
Published: (2024)
Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
by: Liu, Shuhang, et al.
Published: (2025)
by: Liu, Shuhang, et al.
Published: (2025)
Similar Items
-
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024) -
RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction
by: Zhong, Yufeng, et al.
Published: (2025) -
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
by: Xiao, Baihui, et al.
Published: (2025) -
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024) -
RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
by: Zheng, Liming, et al.
Published: (2024)