Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jinkun, Chi, Haohan, Zhang, Lingfeng, Xie, Yifan, Wang, YuAn, Chen, Long, Ye, Hangjun, Hao, Xiaoshuai, Ding, Wenbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
von: Xie, Yifan, et al.
Veröffentlicht: (2026)
von: Xie, Yifan, et al.
Veröffentlicht: (2026)
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2026)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2026)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
von: Fan, Cunxin, et al.
Veröffentlicht: (2025)
von: Fan, Cunxin, et al.
Veröffentlicht: (2025)
Learning to Navigate Socially Through Proactive Risk Perception
von: Xiao, Erjia, et al.
Veröffentlicht: (2025)
von: Xiao, Erjia, et al.
Veröffentlicht: (2025)
Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
Lean-STaR: Learning to Interleave Thinking and Proving
von: Lin, Haohan, et al.
Veröffentlicht: (2024)
von: Lin, Haohan, et al.
Veröffentlicht: (2024)
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
von: Liu, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Liu, Yuanzhe, et al.
Veröffentlicht: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning
von: Zhao, Haokun, et al.
Veröffentlicht: (2026)
von: Zhao, Haokun, et al.
Veröffentlicht: (2026)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
von: Tan, Huajie, et al.
Veröffentlicht: (2026)
von: Tan, Huajie, et al.
Veröffentlicht: (2026)
OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
von: Liu, Yushan, et al.
Veröffentlicht: (2026)
von: Liu, Yushan, et al.
Veröffentlicht: (2026)
A Backbone for Long-Horizon Robot Task Understanding
von: Chen, Xiaoshuai, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoshuai, et al.
Veröffentlicht: (2024)
SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction
von: Fu, Haoxiang, et al.
Veröffentlicht: (2026)
von: Fu, Haoxiang, et al.
Veröffentlicht: (2026)
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Long-Horizon Manipulation via Trace-Conditioned VLA Planning
von: Liu, Isabella, et al.
Veröffentlicht: (2026)
von: Liu, Isabella, et al.
Veröffentlicht: (2026)
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
von: Feng, Yunhai, et al.
Veröffentlicht: (2025)
von: Feng, Yunhai, et al.
Veröffentlicht: (2025)
Weather-Conditioned Branch Routing for Robust LiDAR-Radar 3D Object Detection
von: Li, Hongsheng, et al.
Veröffentlicht: (2026)
von: Li, Hongsheng, et al.
Veröffentlicht: (2026)
Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation
von: Guo, Xinying, et al.
Veröffentlicht: (2026)
von: Guo, Xinying, et al.
Veröffentlicht: (2026)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
von: Li, Yunfei, et al.
Veröffentlicht: (2025)
von: Li, Yunfei, et al.
Veröffentlicht: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
von: Hong, Ilgee, et al.
Veröffentlicht: (2025)
von: Hong, Ilgee, et al.
Veröffentlicht: (2025)
Simple o3: Towards Interleaved Vision-Language Reasoning
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
von: Chen, Zixuan, et al.
Veröffentlicht: (2025)
von: Chen, Zixuan, et al.
Veröffentlicht: (2025)
Learning Bimanual Cloth Manipulation with Vision-based Tactile Sensing via Single Robotic Arm
von: Lee, Dongmyoung, et al.
Veröffentlicht: (2026)
von: Lee, Dongmyoung, et al.
Veröffentlicht: (2026)
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
von: Chen, Qianzhong, et al.
Veröffentlicht: (2025)
von: Chen, Qianzhong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
von: Xie, Yifan, et al.
Veröffentlicht: (2026) -
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025) -
Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2026) -
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
von: Fan, Cunxin, et al.
Veröffentlicht: (2025) -
Learning to Navigate Socially Through Proactive Risk Perception
von: Xiao, Erjia, et al.
Veröffentlicht: (2025)