GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Tong, Yang, Yijun, Xing, Junliang, Shi, Yuanchun, Lu, Zongqing, Ye, Deheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
von: Wei, Tong, et al.
Veröffentlicht: (2025)
von: Wei, Tong, et al.
Veröffentlicht: (2025)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
Visual Grounding for Object-Level Generalization in Reinforcement Learning
von: Jiang, Haobin, et al.
Veröffentlicht: (2024)
von: Jiang, Haobin, et al.
Veröffentlicht: (2024)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
von: Wen, Junwei, et al.
Veröffentlicht: (2026)
von: Wen, Junwei, et al.
Veröffentlicht: (2026)
Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension
von: Li, Lin, et al.
Veröffentlicht: (2025)
von: Li, Lin, et al.
Veröffentlicht: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought
von: Huo, Yu, et al.
Veröffentlicht: (2026)
von: Huo, Yu, et al.
Veröffentlicht: (2026)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
von: Wu, Jiahe, et al.
Veröffentlicht: (2026)
von: Wu, Jiahe, et al.
Veröffentlicht: (2026)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
von: Ouyang, Runqi, et al.
Veröffentlicht: (2025)
von: Ouyang, Runqi, et al.
Veröffentlicht: (2025)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Thought Flow Nets: From Single Predictions to Trains of Model Thought
von: Schuff, Hendrik, et al.
Veröffentlicht: (2021)
von: Schuff, Hendrik, et al.
Veröffentlicht: (2021)
Exploring Reliable PPG Authentication on Smartwatches in Daily Scenarios
von: Tang, Jiankai, et al.
Veröffentlicht: (2025)
von: Tang, Jiankai, et al.
Veröffentlicht: (2025)
Reinforcing Structured Chain-of-Thought for Video Understanding
von: Wang, Peiyao, et al.
Veröffentlicht: (2026)
von: Wang, Peiyao, et al.
Veröffentlicht: (2026)
VQ-Jarvis: Retrieval-Augmented Video Restoration Agent with Sharp Vision and Fast Thought
von: Zhang, Xuanyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xuanyu, et al.
Veröffentlicht: (2026)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
von: Lan, Xing, et al.
Veröffentlicht: (2024)
von: Lan, Xing, et al.
Veröffentlicht: (2024)
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
von: Tang, Bohao, et al.
Veröffentlicht: (2025)
von: Tang, Bohao, et al.
Veröffentlicht: (2025)
$AutoDrive\text{-}P^3$: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-Tuning
von: Ye, Yuqi, et al.
Veröffentlicht: (2026)
von: Ye, Yuqi, et al.
Veröffentlicht: (2026)
Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification
von: Sriram, Ananth, et al.
Veröffentlicht: (2026)
von: Sriram, Ananth, et al.
Veröffentlicht: (2026)
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
von: Xie, Qinghongbing, et al.
Veröffentlicht: (2025)
von: Xie, Qinghongbing, et al.
Veröffentlicht: (2025)
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
von: Wu, Yuan, et al.
Veröffentlicht: (2026)
von: Wu, Yuan, et al.
Veröffentlicht: (2026)
Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
von: Deng, Zekai, et al.
Veröffentlicht: (2025)
von: Deng, Zekai, et al.
Veröffentlicht: (2025)
GoT-CQA: Graph-of-Thought Guided Compositional Reasoning for Chart Question Answering
von: Zhang, Lingling, et al.
Veröffentlicht: (2024)
von: Zhang, Lingling, et al.
Veröffentlicht: (2024)
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
von: Jiang, Yue, et al.
Veröffentlicht: (2024)
von: Jiang, Yue, et al.
Veröffentlicht: (2024)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
von: Ma, Ji, et al.
Veröffentlicht: (2026)
von: Ma, Ji, et al.
Veröffentlicht: (2026)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
Generative Visual Chain-of-Thought for Image Editing
von: Yin, Zijin, et al.
Veröffentlicht: (2026)
von: Yin, Zijin, et al.
Veröffentlicht: (2026)
Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)
von: Liu, Yizhou, et al.
Veröffentlicht: (2025)
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
Fuel Gauge: Estimating Chain-of-Thought Length Ahead of Time in Large Multimodal Models
von: Yang, Yuedong, et al.
Veröffentlicht: (2026)
von: Yang, Yuedong, et al.
Veröffentlicht: (2026)
MaCTG: Multi-Agent Collaborative Thought Graph for Automatic Programming
von: Zhao, Zixiao, et al.
Veröffentlicht: (2024)
von: Zhao, Zixiao, et al.
Veröffentlicht: (2024)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
von: Gao, Timin, et al.
Veröffentlicht: (2024)
von: Gao, Timin, et al.
Veröffentlicht: (2024)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
von: Chen, Wei, et al.
Veröffentlicht: (2026)
von: Chen, Wei, et al.
Veröffentlicht: (2026)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
von: Wei, Tong, et al.
Veröffentlicht: (2025) -
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025) -
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
von: Wei, Yanbin, et al.
Veröffentlicht: (2026) -
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024) -
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
von: Shi, Yudi, et al.
Veröffentlicht: (2024)