RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Zijun, Killick, George, McCreadie, Richard, Camarasa, Gerardo Aragon |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
LaCViT: A Label-aware Contrastive Fine-tuning Framework for Vision Transformers
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
CrisisViT: A Robust Vision Transformer for Crisis Image Classification
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
CLCE: An Approach to Refining Cross-Entropy and Contrastive Learning for Optimized Learning Fusion
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
by: Li, Jianxiong, et al.
Published: (2024)
by: Li, Jianxiong, et al.
Published: (2024)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboEye: Enhancing 2D Robotic Object Identification with Selective 3D Geometric Keypoint Matching
by: Zhang, Xingwu, et al.
Published: (2025)
by: Zhang, Xingwu, et al.
Published: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
by: Wang, Sheng
Published: (2025)
by: Wang, Sheng
Published: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation
by: Zhuang, Lipeng, et al.
Published: (2024)
by: Zhuang, Lipeng, et al.
Published: (2024)
RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis
by: Mu, Yao, et al.
Published: (2024)
by: Mu, Yao, et al.
Published: (2024)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
by: Long, Zijun, et al.
Published: (2025)
by: Long, Zijun, et al.
Published: (2025)
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
by: Han, Songhao, et al.
Published: (2025)
by: Han, Songhao, et al.
Published: (2025)
Is Training Necessary for Anomaly Detection?
by: Zhang, Xingwu, et al.
Published: (2026)
by: Zhang, Xingwu, et al.
Published: (2026)
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
by: Yang, Yuyuan, et al.
Published: (2026)
by: Yang, Yuyuan, et al.
Published: (2026)
RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
by: Mao, Weixin, et al.
Published: (2024)
by: Mao, Weixin, et al.
Published: (2024)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
by: Guruprasad, Pranav, et al.
Published: (2024)
by: Guruprasad, Pranav, et al.
Published: (2024)
RoboPearls: Editable Video Simulation for Robot Manipulation
by: Tang, Tao, et al.
Published: (2025)
by: Tang, Tao, et al.
Published: (2025)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
by: Xu, Peiran, et al.
Published: (2026)
by: Xu, Peiran, et al.
Published: (2026)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
LLM-Grounded Dynamic Task Planning with Hierarchical Temporal Logic for Human-Aware Multi-Robot Collaboration
by: Hu, Shuyuan, et al.
Published: (2026)
by: Hu, Shuyuan, et al.
Published: (2026)
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
by: Huang, Yiqi, et al.
Published: (2025)
by: Huang, Yiqi, et al.
Published: (2025)
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
by: Ji, Yuheng, et al.
Published: (2025)
by: Ji, Yuheng, et al.
Published: (2025)
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
by: Bai, Yu, et al.
Published: (2026)
by: Bai, Yu, et al.
Published: (2026)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
by: Tong, Xinyang, et al.
Published: (2024)
by: Tong, Xinyang, et al.
Published: (2024)
Enabling the Sense of Self in a Dual-Arm Robot
by: AlQallaf, Ali, et al.
Published: (2020)
by: AlQallaf, Ali, et al.
Published: (2020)
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
by: Shao, Rui, et al.
Published: (2025)
by: Shao, Rui, et al.
Published: (2025)
RoboGrasp: A Universal Grasping Policy for Robust Robotic Control
by: Huang, Yiqi, et al.
Published: (2025)
by: Huang, Yiqi, et al.
Published: (2025)
Similar Items
-
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023) -
LaCViT: A Label-aware Contrastive Fine-tuning Framework for Vision Transformers
by: Long, Zijun, et al.
Published: (2023) -
Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning
by: Long, Zijun, et al.
Published: (2024) -
CrisisViT: A Robust Vision Transformer for Crisis Image Classification
by: Long, Zijun, et al.
Published: (2024) -
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)