Saved in:
| Main Authors: | Luan, Zhirong, Lai, Yujun, Huang, Rundong, Lan, Xiaruiqi, Chen, Liangjun, Chen, Badong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.03699 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hierarchical Large Language Models in Cloud Edge End Architecture for Heterogeneous Robot Cluster Control
by: Luan, Zhirong, et al.
Published: (2024)
by: Luan, Zhirong, et al.
Published: (2024)
LIT: Large Language Model Driven Intention Tracking for Proactive Human-Robot Collaboration -- A Robot Sous-Chef Application
by: Huang, Zhe, et al.
Published: (2024)
by: Huang, Zhe, et al.
Published: (2024)
Soft Prompt Generation for Domain Generalization
by: Bai, Shuanghao, et al.
Published: (2024)
by: Bai, Shuanghao, et al.
Published: (2024)
S3E: A Multi-Robot Multimodal Dataset for Collaborative SLAM
by: Feng, Dapeng, et al.
Published: (2022)
by: Feng, Dapeng, et al.
Published: (2022)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
by: Shao, Rui, et al.
Published: (2025)
by: Shao, Rui, et al.
Published: (2025)
Patch-Based Spatial Authorship Attribution in Human-Robot Collaborative Paintings
by: Chen, Eric, et al.
Published: (2026)
by: Chen, Eric, et al.
Published: (2026)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
by: Tong, Xinyang, et al.
Published: (2024)
by: Tong, Xinyang, et al.
Published: (2024)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
Causal World Modeling for Robot Control
by: Li, Lin, et al.
Published: (2026)
by: Li, Lin, et al.
Published: (2026)
Dual-Path Stable Soft Prompt Generation for Domain Generalization
by: Zhang, Yuedi, et al.
Published: (2025)
by: Zhang, Yuedi, et al.
Published: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
by: Liu, Kangcheng, et al.
Published: (2023)
by: Liu, Kangcheng, et al.
Published: (2023)
Swarm-SLAM : Sparse Decentralized Collaborative Simultaneous Localization and Mapping Framework for Multi-Robot Systems
by: Lajoie, Pierre-Yves, et al.
Published: (2023)
by: Lajoie, Pierre-Yves, et al.
Published: (2023)
Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy Prediction
by: Chen, Cheng, et al.
Published: (2025)
by: Chen, Cheng, et al.
Published: (2025)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
Prompt-based Distribution Alignment for Unsupervised Domain Adaptation
by: Bai, Shuanghao, et al.
Published: (2023)
by: Bai, Shuanghao, et al.
Published: (2023)
Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities
by: Zhou, Yiyun, et al.
Published: (2025)
by: Zhou, Yiyun, et al.
Published: (2025)
Robot Manipulation in Salient Vision through Referring Image Segmentation and Geometric Constraints
by: Jiang, Chen, et al.
Published: (2024)
by: Jiang, Chen, et al.
Published: (2024)
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
by: Han, Songhao, et al.
Published: (2025)
by: Han, Songhao, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
A Survey on Improving Human Robot Collaboration through Vision-and-Language Navigation
by: Yakolli, Nivedan, et al.
Published: (2025)
by: Yakolli, Nivedan, et al.
Published: (2025)
Multi-Camera Hand-Eye Calibration for Human-Robot Collaboration in Industrial Robotic Workcells
by: Allegro, Davide, et al.
Published: (2024)
by: Allegro, Davide, et al.
Published: (2024)
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
by: Chen, Yuan, et al.
Published: (2024)
by: Chen, Yuan, et al.
Published: (2024)
VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
PhysPart: Physically Plausible Part Completion for Interactable Objects
by: Luo, Rundong, et al.
Published: (2024)
by: Luo, Rundong, et al.
Published: (2024)
Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees
by: Thumm, Jakob, et al.
Published: (2026)
by: Thumm, Jakob, et al.
Published: (2026)
Evaluating Pointing Gestures for Target Selection in Human-Robot Collaboration
by: Sassali, Noora, et al.
Published: (2025)
by: Sassali, Noora, et al.
Published: (2025)
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026)
by: Liu, Mengzhen, et al.
Published: (2026)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts
by: Yang, Yanting, et al.
Published: (2024)
by: Yang, Yanting, et al.
Published: (2024)
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
by: Mei, Haiyang, et al.
Published: (2025)
by: Mei, Haiyang, et al.
Published: (2025)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
Similar Items
-
Hierarchical Large Language Models in Cloud Edge End Architecture for Heterogeneous Robot Cluster Control
by: Luan, Zhirong, et al.
Published: (2024) -
LIT: Large Language Model Driven Intention Tracking for Proactive Human-Robot Collaboration -- A Robot Sous-Chef Application
by: Huang, Zhe, et al.
Published: (2024) -
Soft Prompt Generation for Domain Generalization
by: Bai, Shuanghao, et al.
Published: (2024) -
S3E: A Multi-Robot Multimodal Dataset for Collaborative SLAM
by: Feng, Dapeng, et al.
Published: (2022) -
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)