Octopus: Embodied Vision-Language Programmer from Environmental Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jingkang, Dong, Yuhao, Liu, Shuai, Li, Bo, Wang, Ziyue, Jiang, Chencheng, Tan, Haoran, Kang, Jiamu, Zhang, Yuanhan, Zhou, Kaiyang, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Long Context Transfer from Language to Vision
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
Survey of Vision-Language-Action Models for Embodied Manipulation
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
von: Yu, En, et al.
Veröffentlicht: (2026)
von: Yu, En, et al.
Veröffentlicht: (2026)
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
von: Zhang, Hanxin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanxin, et al.
Veröffentlicht: (2026)
Generalized Out-of-Distribution Detection: A Survey
von: Yang, Jingkang, et al.
Veröffentlicht: (2021)
von: Yang, Jingkang, et al.
Veröffentlicht: (2021)
PM-Nav: Priori-Map Guided Embodied Navigation in Functional Buildings
von: Gao, Jiang, et al.
Veröffentlicht: (2026)
von: Gao, Jiang, et al.
Veröffentlicht: (2026)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
VLP: Vision-Language Preference Learning for Embodied Manipulation
von: Liu, Runze, et al.
Veröffentlicht: (2025)
von: Liu, Runze, et al.
Veröffentlicht: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
von: Li, Boyu, et al.
Veröffentlicht: (2026)
von: Li, Boyu, et al.
Veröffentlicht: (2026)
Embodied Scene Understanding for Vision Language Models via MetaVQA
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
von: Yang, Zhejian, et al.
Veröffentlicht: (2025)
von: Yang, Zhejian, et al.
Veröffentlicht: (2025)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
von: Zhang, Jiyao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiyao, et al.
Veröffentlicht: (2026)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
von: Xu, Haiweng, et al.
Veröffentlicht: (2026)
von: Xu, Haiweng, et al.
Veröffentlicht: (2026)
NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
von: Windecker, Tim, et al.
Veröffentlicht: (2025)
von: Windecker, Tim, et al.
Veröffentlicht: (2025)
Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models
von: Soedarmadji, Saraswati, et al.
Veröffentlicht: (2025)
von: Soedarmadji, Saraswati, et al.
Veröffentlicht: (2025)
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
von: Tai, Cong, et al.
Veröffentlicht: (2025)
von: Tai, Cong, et al.
Veröffentlicht: (2025)
A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration
von: Xu, Kuan, et al.
Veröffentlicht: (2026)
von: Xu, Kuan, et al.
Veröffentlicht: (2026)
Stable Language Guidance for Vision-Language-Action Models
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
von: Bo, Zitong, et al.
Veröffentlicht: (2025)
von: Bo, Zitong, et al.
Veröffentlicht: (2025)
Embodied Navigation Foundation Model
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2025)
ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making
von: Dai, Liu, et al.
Veröffentlicht: (2025)
von: Dai, Liu, et al.
Veröffentlicht: (2025)
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025)
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
von: Chen, William, et al.
Veröffentlicht: (2026)
von: Chen, William, et al.
Veröffentlicht: (2026)
Distributed and Consistent Multi-Robot Visual-Inertial-Ranging Odometry on Lie Groups
von: Kang, Ziwei, et al.
Veröffentlicht: (2026)
von: Kang, Ziwei, et al.
Veröffentlicht: (2026)
Lifelong Embodied Navigation Learning
von: Wang, Xudong, et al.
Veröffentlicht: (2026)
von: Wang, Xudong, et al.
Veröffentlicht: (2026)
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
SpaceOctopus: An Octopus-inspired Motion Planning Framework for Multi-arm Space Robot
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback
von: Shu, Junyang, et al.
Veröffentlicht: (2025)
von: Shu, Junyang, et al.
Veröffentlicht: (2025)
HMR-1: Hierarchical Massage Robot with Vision-Language-Model for Embodied Healthcare
von: Xu, Rongtao, et al.
Veröffentlicht: (2026)
von: Xu, Rongtao, et al.
Veröffentlicht: (2026)
Embodied Hazard Mitigation using Vision-Language Models for Autonomous Mobile Robots
von: Sotomi, Oluwadamilola, et al.
Veröffentlicht: (2025)
von: Sotomi, Oluwadamilola, et al.
Veröffentlicht: (2025)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
AIR-Embodied: An Efficient Active 3DGS-based Interaction and Reconstruction Framework with Embodied Large Language Model
von: Qi, Zhenghao, et al.
Veröffentlicht: (2024)
von: Qi, Zhenghao, et al.
Veröffentlicht: (2024)
PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
von: Ding, Wenbin, et al.
Veröffentlicht: (2025)
von: Ding, Wenbin, et al.
Veröffentlicht: (2025)
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
von: Zhao, Xinxin, et al.
Veröffentlicht: (2024)
von: Zhao, Xinxin, et al.
Veröffentlicht: (2024)
ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
von: Wang, Teng, et al.
Veröffentlicht: (2025)
von: Wang, Teng, et al.
Veröffentlicht: (2025)
TravExplorer: Cross-Floor Embodied Exploration via Traversability-Aware 3-D Planning
von: Zheng, Han, et al.
Veröffentlicht: (2026)
von: Zheng, Han, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Long Context Transfer from Language to Vision
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024) -
Survey of Vision-Language-Action Models for Embodied Manipulation
von: Li, Haoran, et al.
Veröffentlicht: (2025) -
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026) -
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
von: Yu, En, et al.
Veröffentlicht: (2026) -
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
von: Zhang, Hanxin, et al.
Veröffentlicht: (2026)