Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Dong-Hee, Tan, Reuben, Kim, Donghyun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
by: Kim, Dong-Hee, et al.
Published: (2025)
by: Kim, Dong-Hee, et al.
Published: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026)
by: Lim, Byeonggeuk, et al.
Published: (2026)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023)
by: Kim, Sungyeon, et al.
Published: (2023)
Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2026)
by: Kim, Donghyun, et al.
Published: (2026)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025)
by: Tian, Shulin, et al.
Published: (2025)
Rethinking Chain-of-Thought Reasoning for Videos
by: Zhong, Yiwu, et al.
Published: (2025)
by: Zhong, Yiwu, et al.
Published: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
by: Hwang, Hochul, et al.
Published: (2024)
by: Hwang, Hochul, et al.
Published: (2024)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
by: Guo, Guangfu, et al.
Published: (2026)
by: Guo, Guangfu, et al.
Published: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
by: Jain, Riddhi, et al.
Published: (2025)
by: Jain, Riddhi, et al.
Published: (2025)
Rethinking Visual Information Processing in Multimodal LLMs
by: Kim, Dongwan, et al.
Published: (2025)
by: Kim, Dongwan, et al.
Published: (2025)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
by: Yi, Shixin, et al.
Published: (2025)
by: Yi, Shixin, et al.
Published: (2025)
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
by: Corbière, Charles, et al.
Published: (2025)
by: Corbière, Charles, et al.
Published: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Visual Accommodation: Rethinking Image Scale as a Learnable Variable for Object Detection
by: Seo, Daeun, et al.
Published: (2024)
by: Seo, Daeun, et al.
Published: (2024)
VisAgent: Narrative-Preserving Story Visualization Framework
by: Kim, Seungkwon, et al.
Published: (2025)
by: Kim, Seungkwon, et al.
Published: (2025)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025)
by: An, Sojung, et al.
Published: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2024)
by: Kim, Kibum, et al.
Published: (2024)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
by: Lee, Eunsoo, et al.
Published: (2026)
by: Lee, Eunsoo, et al.
Published: (2026)
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
by: Lu, Wenting, et al.
Published: (2026)
by: Lu, Wenting, et al.
Published: (2026)
Training-Free Label Space Alignment for Universal Domain Adaptation
by: Lee, Dujin, et al.
Published: (2025)
by: Lee, Dujin, et al.
Published: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
by: Wu, Yuan, et al.
Published: (2026)
by: Wu, Yuan, et al.
Published: (2026)
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
by: Guo, Yuchen, et al.
Published: (2026)
by: Guo, Yuchen, et al.
Published: (2026)
Reinforcing Structured Chain-of-Thought for Video Understanding
by: Wang, Peiyao, et al.
Published: (2026)
by: Wang, Peiyao, et al.
Published: (2026)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025)
by: Kim, Hee-Seon, et al.
Published: (2025)
ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language Model
by: Chen, Xiaoshu, et al.
Published: (2026)
by: Chen, Xiaoshu, et al.
Published: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
Controllable Navigation Instruction Generation with Chain of Thought Prompting
by: Kong, Xianghao, et al.
Published: (2024)
by: Kong, Xianghao, et al.
Published: (2024)
PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D Data
by: Yang, ChangHee, et al.
Published: (2025)
by: Yang, ChangHee, et al.
Published: (2025)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
by: He, Hulingxiao, et al.
Published: (2026)
by: He, Hulingxiao, et al.
Published: (2026)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Similar Items
-
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
by: Kim, Dong-Hee, et al.
Published: (2025) -
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026) -
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023) -
Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2026) -
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025)