Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Cunxin, Jia, Xiaosong, Sun, Yihang, Wang, Yixiao, Wei, Jianglan, Gong, Ziyang, Zhao, Xiangyu, Tomizuka, Masayoshi, Yang, Xue, Yan, Junchi, Ding, Mingyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
von: Tang, Weiliang, et al.
Veröffentlicht: (2025)
von: Tang, Weiliang, et al.
Veröffentlicht: (2025)
DexH2R: Task-oriented Dexterous Manipulation from Human to Robots
von: Zhao, Shuqi, et al.
Veröffentlicht: (2024)
von: Zhao, Shuqi, et al.
Veröffentlicht: (2024)
Physics-Aware Robotic Palletization with Online Masking Inference
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
Nonparametric Inverse Dynamic Models for Multimodal Interactive Robots
von: Haninger, Kevin, et al.
Veröffentlicht: (2019)
von: Haninger, Kevin, et al.
Veröffentlicht: (2019)
Make Your VLA More Robust Without More Data By Interleaving Motion Planning
von: Choe, Dan BW, et al.
Veröffentlicht: (2026)
von: Choe, Dan BW, et al.
Veröffentlicht: (2026)
TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control
von: Sun, Yuteng, et al.
Veröffentlicht: (2026)
von: Sun, Yuteng, et al.
Veröffentlicht: (2026)
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving
von: Liu, Qiqi, et al.
Veröffentlicht: (2026)
von: Liu, Qiqi, et al.
Veröffentlicht: (2026)
DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulation
von: Liang, Zhixuan, et al.
Veröffentlicht: (2024)
von: Liang, Zhixuan, et al.
Veröffentlicht: (2024)
Robust In-Hand Manipulation with Extrinsic Contacts
von: Liang, Boyuan, et al.
Veröffentlicht: (2024)
von: Liang, Boyuan, et al.
Veröffentlicht: (2024)
WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
Unified Manipulability and Compliance Analysis of Modular Soft-Rigid Hybrid Fingers
von: Zhou, Jianshu, et al.
Veröffentlicht: (2025)
von: Zhou, Jianshu, et al.
Veröffentlicht: (2025)
Adaptive Linear Path Model-Based Diffusion
von: Shimizu, Yutaka, et al.
Veröffentlicht: (2026)
von: Shimizu, Yutaka, et al.
Veröffentlicht: (2026)
Leveraging Extrinsic Dexterity for Occluded Grasping on Grasp Constraining Walls
von: Kobashi, Keita, et al.
Veröffentlicht: (2025)
von: Kobashi, Keita, et al.
Veröffentlicht: (2025)
Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning
von: Wang, Yixiao, et al.
Veröffentlicht: (2024)
von: Wang, Yixiao, et al.
Veröffentlicht: (2024)
P2 Explore: Efficient Exploration in Unknown Cluttered Environment with Floor Plan Prediction
von: Song, Kun, et al.
Veröffentlicht: (2024)
von: Song, Kun, et al.
Veröffentlicht: (2024)
Prismatic-Bending Transformable (PBT) Joint for a Modular, Foldable Manipulator with Enhanced Reachability and Dexterity
von: Zhou, Jianshu, et al.
Veröffentlicht: (2025)
von: Zhou, Jianshu, et al.
Veröffentlicht: (2025)
Harnessing with Twisting: Single-Arm Deformable Linear Object Manipulation for Industrial Harnessing Task
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
Fractional-order Modeling for Nonlinear Soft Actuators via Particle Swarm Optimization
von: Yang, Wu-Te, et al.
Veröffentlicht: (2025)
von: Yang, Wu-Te, et al.
Veröffentlicht: (2025)
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
von: Wang, Yixiao, et al.
Veröffentlicht: (2025)
von: Wang, Yixiao, et al.
Veröffentlicht: (2025)
DrPlanner: Diagnosis and Repair of Motion Planners for Automated Vehicles Using Large Language Models
von: Lin, Yuanfei, et al.
Veröffentlicht: (2024)
von: Lin, Yuanfei, et al.
Veröffentlicht: (2024)
Joint Pedestrian Trajectory Prediction through Posterior Sampling
von: Lin, Haotian, et al.
Veröffentlicht: (2024)
von: Lin, Haotian, et al.
Veröffentlicht: (2024)
Robust Model-Based In-Hand Manipulation with Integrated Real-Time Motion-Contact Planning and Tracking
von: Jiang, Yongpeng, et al.
Veröffentlicht: (2025)
von: Jiang, Yongpeng, et al.
Veröffentlicht: (2025)
Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach
von: Jiang, Yongpeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yongpeng, et al.
Veröffentlicht: (2024)
Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
Language-Driven Policy Distillation for Cooperative Driving in Multi-Agent Reinforcement Learning
von: Liu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2024)
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
von: Jia, Xiaosong, et al.
Veröffentlicht: (2026)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2026)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
Bridging the Sim-to-Real Gap with Dynamic Compliance Tuning for Industrial Insertion
von: Zhang, Xiang, et al.
Veröffentlicht: (2023)
von: Zhang, Xiang, et al.
Veröffentlicht: (2023)
CLAW: Composable Language-Annotated Whole-body Motion Generation
von: Cao, Jianuo, et al.
Veröffentlicht: (2026)
von: Cao, Jianuo, et al.
Veröffentlicht: (2026)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
von: Luo, Yuankai, et al.
Veröffentlicht: (2026)
von: Luo, Yuankai, et al.
Veröffentlicht: (2026)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task Execution
von: Liang, Zhixuan, et al.
Veröffentlicht: (2023)
von: Liang, Zhixuan, et al.
Veröffentlicht: (2023)
Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2)
von: Li, Qifeng, et al.
Veröffentlicht: (2024)
von: Li, Qifeng, et al.
Veröffentlicht: (2024)
Underactuated Control of Multiple Soft Pneumatic Actuators via Stable Inversion
von: Yang, Wu-Te, et al.
Veröffentlicht: (2024)
von: Yang, Wu-Te, et al.
Veröffentlicht: (2024)
Optimized Design of a Soft Actuator Considering Force/Torque, Bendability, and Controllability via an Approximated Structure
von: Yang, Wu-Te, et al.
Veröffentlicht: (2023)
von: Yang, Wu-Te, et al.
Veröffentlicht: (2023)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
von: Liu, Jinkun, et al.
Veröffentlicht: (2026) -
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
von: Hu, Yucheng, et al.
Veröffentlicht: (2026) -
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
von: Tang, Weiliang, et al.
Veröffentlicht: (2025) -
DexH2R: Task-oriented Dexterous Manipulation from Human to Robots
von: Zhao, Shuqi, et al.
Veröffentlicht: (2024) -
Physics-Aware Robotic Palletization with Online Masking Inference
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)