VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lim, Byeonggeuk, Kim, Kyeonghyun, Yun, JungMin, Kim, YoungBin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
Query, Decompose, Compress: Structured Query Expansion for Efficient Multi-Hop Retrieval
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
von: Yun, JungMin, et al.
Veröffentlicht: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
von: Choi, Hojun, et al.
Veröffentlicht: (2025)
von: Choi, Hojun, et al.
Veröffentlicht: (2025)
Colorful Cutout: Enhancing Image Data Augmentation with Curriculum Learning
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025)
von: Kim, Junho, et al.
Veröffentlicht: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
von: Acuna, David, et al.
Veröffentlicht: (2025)
von: Acuna, David, et al.
Veröffentlicht: (2025)
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
von: Ng, Chee, et al.
Veröffentlicht: (2025)
von: Ng, Chee, et al.
Veröffentlicht: (2025)
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
von: Yi, Shixin, et al.
Veröffentlicht: (2025)
von: Yi, Shixin, et al.
Veröffentlicht: (2025)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)
von: Bai, Sule, et al.
Veröffentlicht: (2025)
Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior
von: Wang, Sheng, et al.
Veröffentlicht: (2025)
von: Wang, Sheng, et al.
Veröffentlicht: (2025)
GUIDE-CoT: Goal-driven and User-Informed Dynamic Estimation for Pedestrian Trajectory using Chain-of-Thought
von: Kim, Sungsik, et al.
Veröffentlicht: (2025)
von: Kim, Sungsik, et al.
Veröffentlicht: (2025)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
von: Fang, Hao, et al.
Veröffentlicht: (2026)
von: Fang, Hao, et al.
Veröffentlicht: (2026)
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
von: Qi, Yu, et al.
Veröffentlicht: (2025)
von: Qi, Yu, et al.
Veröffentlicht: (2025)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)
PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
von: Zhong, Chunlin, et al.
Veröffentlicht: (2025)
von: Zhong, Chunlin, et al.
Veröffentlicht: (2025)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026) -
Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026) -
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026) -
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
von: Choi, Juhwan, et al.
Veröffentlicht: (2024) -
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)