Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Sunil, Zhao, Bowen, Dirac, Leo, Varshavskaya, Paulina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Fine-tuning Vision Classifiers On A Budget
by: Kumar, Sunil, et al.
Published: (2024)
by: Kumar, Sunil, et al.
Published: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
by: Xia, Canming, et al.
Published: (2026)
by: Xia, Canming, et al.
Published: (2026)
Diffusion for World Modeling: Visual Details Matter in Atari
by: Alonso, Eloi, et al.
Published: (2024)
by: Alonso, Eloi, et al.
Published: (2024)
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
by: Cho, Jang Hyun, et al.
Published: (2025)
by: Cho, Jang Hyun, et al.
Published: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization
by: Karim, Mohammed Asad, et al.
Published: (2026)
by: Karim, Mohammed Asad, et al.
Published: (2026)
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024)
by: Cooper, Avi, et al.
Published: (2024)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
Learning Visual Abstract Reasoning through Dual-Stream Networks
by: Zhao, Kai, et al.
Published: (2024)
by: Zhao, Kai, et al.
Published: (2024)
CARES: Context-Aware Resolution Selector for VLMs
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
DASH: Detection and Assessment of Systematic Hallucinations of VLMs
by: Augustin, Maximilian, et al.
Published: (2025)
by: Augustin, Maximilian, et al.
Published: (2025)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)
by: Moon, Jihyun, et al.
Published: (2025)
Evaluation of Test-Time Adaptation Under Computational Time Constraints
by: Alfarra, Motasem, et al.
Published: (2023)
by: Alfarra, Motasem, et al.
Published: (2023)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
by: Ding, Xi, et al.
Published: (2024)
by: Ding, Xi, et al.
Published: (2024)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
by: Karamcheti, Siddharth, et al.
Published: (2024)
by: Karamcheti, Siddharth, et al.
Published: (2024)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
by: Gavrikov, Paul, et al.
Published: (2025)
by: Gavrikov, Paul, et al.
Published: (2025)
Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Do We Need Large VLMs for Spotting Soccer Actions?
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)
by: Jaiswal, Shantanu, et al.
Published: (2024)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
Developing a Resource-Constraint EdgeAI model for Surface Defect Detection
by: Mih, Atah Nuh, et al.
Published: (2023)
by: Mih, Atah Nuh, et al.
Published: (2023)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
by: Liu, Li, et al.
Published: (2024)
by: Liu, Li, et al.
Published: (2024)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
by: Reddy, N Dinesh, et al.
Published: (2025)
by: Reddy, N Dinesh, et al.
Published: (2025)
PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
by: Meng, Ziqiao, et al.
Published: (2025)
by: Meng, Ziqiao, et al.
Published: (2025)
SPHINX: A Synthetic Environment for Visual Perception and Reasoning
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
Similar Items
-
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
by: Zhao, Bowen, et al.
Published: (2024) -
Fine-tuning Vision Classifiers On A Budget
by: Kumar, Sunil, et al.
Published: (2024) -
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025) -
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026) -
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)