Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Kangcheng, Liu, Yong-Jin, Chen, Baoquan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
von: Wen, Xin, et al.
Veröffentlicht: (2025)
von: Wen, Xin, et al.
Veröffentlicht: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
von: Shao, Rui, et al.
Veröffentlicht: (2025)
von: Shao, Rui, et al.
Veröffentlicht: (2025)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities
von: Zhou, Yiyun, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyun, et al.
Veröffentlicht: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
A Touch, Vision, and Language Dataset for Multimodal Alignment
von: Fu, Letian, et al.
Veröffentlicht: (2024)
von: Fu, Letian, et al.
Veröffentlicht: (2024)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
Continuous Object State Recognition for Cooking Robots Using Pre-Trained Vision-Language Models and Black-box Optimization
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2024)
Synthetic Dataset Generation for Autonomous Mobile Robots Using 3D Gaussian Splatting for Vision Training
von: Deogan, Aneesh, et al.
Veröffentlicht: (2025)
von: Deogan, Aneesh, et al.
Veröffentlicht: (2025)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
von: Liu, Chuhang, et al.
Veröffentlicht: (2026)
von: Liu, Chuhang, et al.
Veröffentlicht: (2026)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
von: Din, Muhayy Ud, et al.
Veröffentlicht: (2025)
von: Din, Muhayy Ud, et al.
Veröffentlicht: (2025)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
von: Dai, Tingjun, et al.
Veröffentlicht: (2026)
von: Dai, Tingjun, et al.
Veröffentlicht: (2026)
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
von: Rowe, Luke, et al.
Veröffentlicht: (2025)
von: Rowe, Luke, et al.
Veröffentlicht: (2025)
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
von: Lu, Yangxiao, et al.
Veröffentlicht: (2024)
von: Lu, Yangxiao, et al.
Veröffentlicht: (2024)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
Zero-Shot 3D Visual Grounding from Vision-Language Models
von: Li, Rong, et al.
Veröffentlicht: (2025)
von: Li, Rong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
von: Wen, Xin, et al.
Veröffentlicht: (2025) -
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024) -
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025) -
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025) -
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)