RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
Fuente:
arXiv
Salvato in:
| Autore principale: | Wang, Sheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
di: Yuan, Wentao, et al.
Pubblicazione: (2024)
di: Yuan, Wentao, et al.
Pubblicazione: (2024)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
di: Fan, Qingyu, et al.
Pubblicazione: (2026)
di: Fan, Qingyu, et al.
Pubblicazione: (2026)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
di: Wang, Yi Ru, et al.
Pubblicazione: (2025)
di: Wang, Yi Ru, et al.
Pubblicazione: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
di: Gao, Jensen, et al.
Pubblicazione: (2023)
di: Gao, Jensen, et al.
Pubblicazione: (2023)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
di: Kirchner, Sven, et al.
Pubblicazione: (2025)
di: Kirchner, Sven, et al.
Pubblicazione: (2025)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
di: Wang, Boyang, et al.
Pubblicazione: (2026)
di: Wang, Boyang, et al.
Pubblicazione: (2026)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
di: Song, Chan Hee, et al.
Pubblicazione: (2024)
di: Song, Chan Hee, et al.
Pubblicazione: (2024)
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
di: Wei, Ziyu, et al.
Pubblicazione: (2026)
di: Wei, Ziyu, et al.
Pubblicazione: (2026)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
di: Chen, Shizhe, et al.
Pubblicazione: (2025)
di: Chen, Shizhe, et al.
Pubblicazione: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
di: Li, Xin, et al.
Pubblicazione: (2024)
di: Li, Xin, et al.
Pubblicazione: (2024)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
di: Liu, Minghuan, et al.
Pubblicazione: (2025)
di: Liu, Minghuan, et al.
Pubblicazione: (2025)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
di: Han, Yi, et al.
Pubblicazione: (2025)
di: Han, Yi, et al.
Pubblicazione: (2025)
RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation
di: Jiang, Hanxiao, et al.
Pubblicazione: (2024)
di: Jiang, Hanxiao, et al.
Pubblicazione: (2024)
RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis
di: Mu, Yao, et al.
Pubblicazione: (2024)
di: Mu, Yao, et al.
Pubblicazione: (2024)
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
di: Li, Zaijing, et al.
Pubblicazione: (2026)
di: Li, Zaijing, et al.
Pubblicazione: (2026)
RoboEye: Enhancing 2D Robotic Object Identification with Selective 3D Geometric Keypoint Matching
di: Zhang, Xingwu, et al.
Pubblicazione: (2025)
di: Zhang, Xingwu, et al.
Pubblicazione: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
di: Kim, Seungku, et al.
Pubblicazione: (2026)
di: Kim, Seungku, et al.
Pubblicazione: (2026)
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
di: Lin, Xuewu, et al.
Pubblicazione: (2025)
di: Lin, Xuewu, et al.
Pubblicazione: (2025)
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
di: Chen, Zezhou, et al.
Pubblicazione: (2025)
di: Chen, Zezhou, et al.
Pubblicazione: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
di: Polizzi, Vincenzo, et al.
Pubblicazione: (2026)
di: Polizzi, Vincenzo, et al.
Pubblicazione: (2026)
RPMArt: Towards Robust Perception and Manipulation for Articulated Objects
di: Wang, Junbo, et al.
Pubblicazione: (2024)
di: Wang, Junbo, et al.
Pubblicazione: (2024)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
di: Yang, Sheng, et al.
Pubblicazione: (2025)
di: Yang, Sheng, et al.
Pubblicazione: (2025)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
di: Cui, Wenbo, et al.
Pubblicazione: (2025)
di: Cui, Wenbo, et al.
Pubblicazione: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
di: Wen, Yuqing, et al.
Pubblicazione: (2025)
di: Wen, Yuqing, et al.
Pubblicazione: (2025)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
di: Kamboj, Abhi, et al.
Pubblicazione: (2024)
di: Kamboj, Abhi, et al.
Pubblicazione: (2024)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
di: Li, Qixiu, et al.
Pubblicazione: (2025)
di: Li, Qixiu, et al.
Pubblicazione: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
di: Sarowar, Md Selim, et al.
Pubblicazione: (2026)
di: Sarowar, Md Selim, et al.
Pubblicazione: (2026)
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
IRASim: A Fine-Grained World Model for Robot Manipulation
di: Zhu, Fangqi, et al.
Pubblicazione: (2024)
di: Zhu, Fangqi, et al.
Pubblicazione: (2024)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
ClearDepth: Enhanced Stereo Perception of Transparent Objects for Robotic Manipulation
di: Bai, Kaixin, et al.
Pubblicazione: (2024)
di: Bai, Kaixin, et al.
Pubblicazione: (2024)
Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
di: Liu, Yijun, et al.
Pubblicazione: (2025)
di: Liu, Yijun, et al.
Pubblicazione: (2025)
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
di: Wu, Yiming, et al.
Pubblicazione: (2025)
di: Wu, Yiming, et al.
Pubblicazione: (2025)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
di: Lohner, Aaron, et al.
Pubblicazione: (2024)
di: Lohner, Aaron, et al.
Pubblicazione: (2024)
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
di: Rasouli, Amir, et al.
Pubblicazione: (2025)
di: Rasouli, Amir, et al.
Pubblicazione: (2025)
HomeRobot: Open-Vocabulary Mobile Manipulation
di: Yenamandra, Sriram, et al.
Pubblicazione: (2023)
di: Yenamandra, Sriram, et al.
Pubblicazione: (2023)
Documenti analoghi
-
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025) -
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
di: Yuan, Wentao, et al.
Pubblicazione: (2024) -
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
di: Fan, Qingyu, et al.
Pubblicazione: (2026) -
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
di: Wang, Yi Ru, et al.
Pubblicazione: (2025) -
Physically Grounded Vision-Language Models for Robotic Manipulation
di: Gao, Jensen, et al.
Pubblicazione: (2023)