CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Qingqing, Lu, Yao, Kim, Moo Jin, Fu, Zipeng, Zhang, Zhuoyang, Wu, Yecheng, Li, Zhaoshuo, Ma, Qianli, Han, Song, Finn, Chelsea, Handa, Ankur, Liu, Ming-Yu, Xiang, Donglai, Wetzstein, Gordon, Lin, Tsung-Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HumanPlus: Humanoid Shadowing and Imitation from Humans
von: Fu, Zipeng, et al.
Veröffentlicht: (2024)
von: Fu, Zipeng, et al.
Veröffentlicht: (2024)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
von: Mondal, Debjyoti, et al.
Veröffentlicht: (2024)
von: Mondal, Debjyoti, et al.
Veröffentlicht: (2024)
CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning
von: Fang, Yuanheng, et al.
Veröffentlicht: (2025)
von: Fang, Yuanheng, et al.
Veröffentlicht: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
Co-CoT: A Prompt-Based Framework for Collaborative Chain-of-Thought Reasoning
von: Yoo, Seunghyun
Veröffentlicht: (2025)
von: Yoo, Seunghyun
Veröffentlicht: (2025)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
von: Fan, Lin, et al.
Veröffentlicht: (2026)
von: Fan, Lin, et al.
Veröffentlicht: (2026)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
SIM-CoT: Supervised Implicit Chain-of-Thought
von: Wei, Xilin, et al.
Veröffentlicht: (2025)
von: Wei, Xilin, et al.
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
Robotic Control via Embodied Chain-of-Thought Reasoning
von: Zawalski, Michał, et al.
Veröffentlicht: (2024)
von: Zawalski, Michał, et al.
Veröffentlicht: (2024)
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
von: Chen, Shuxu, et al.
Veröffentlicht: (2026)
von: Chen, Shuxu, et al.
Veröffentlicht: (2026)
Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models
von: Das, Arghyadeep, et al.
Veröffentlicht: (2026)
von: Das, Arghyadeep, et al.
Veröffentlicht: (2026)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
GTR-CoT: Graph Traversal as Visual Chain of Thought for Molecular Structure Recognition
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
von: Zheng, Tianshi, et al.
Veröffentlicht: (2025)
CoT-Valve: Length-Compressible Chain-of-Thought Tuning
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
OpenVLA: An Open-Source Vision-Language-Action Model
von: Kim, Moo Jin, et al.
Veröffentlicht: (2024)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
von: Ng, Chee, et al.
Veröffentlicht: (2025)
von: Ng, Chee, et al.
Veröffentlicht: (2025)
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
von: Arnav, Benjamin, et al.
Veröffentlicht: (2025)
von: Arnav, Benjamin, et al.
Veröffentlicht: (2025)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
von: Qi, Yu, et al.
Veröffentlicht: (2025)
von: Qi, Yu, et al.
Veröffentlicht: (2025)
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
von: Jang, Dongjun, et al.
Veröffentlicht: (2025)
von: Jang, Dongjun, et al.
Veröffentlicht: (2025)
Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
von: Tiwari, Nitya, et al.
Veröffentlicht: (2025)
von: Tiwari, Nitya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HumanPlus: Humanoid Shadowing and Imitation from Humans
von: Fu, Zipeng, et al.
Veröffentlicht: (2024) -
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025) -
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2026) -
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2025) -
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)