Saved in:
Bibliographic Details
Main Authors: Hou, Xinhai, Xu, Shaoyuan, Biyani, Manan, Li, Moyan, Liu, Jia, Hollon, Todd C., Wang, Bryan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.19661
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915826712444928
author Hou, Xinhai
Xu, Shaoyuan
Biyani, Manan
Li, Moyan
Liu, Jia
Hollon, Todd C.
Wang, Bryan
author_facet Hou, Xinhai
Xu, Shaoyuan
Biyani, Manan
Li, Moyan
Liu, Jia
Hollon, Todd C.
Wang, Bryan
contents Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful visual reasoning: models may invoke tools on irrelevant regions or ignore tool outputs entirely, yet still guess the correct answer. In this work, we first propose a faithfulness evaluation protocol that measures whether intermediate visual tool outputs (e.g., crops) actually contain the queried evidence. This reveals that recent visual agents achieve high final-answer accuracy but exhibit low rates of faithful tool-use on visual search benchmarks. We then introduce CodeV, a code-based visual agent trained with Tool-Aware Policy Optimization (TAPO). TAPO is a process-level RL framework that augments GRPO with dense rewards defined directly on visual tool inputs and outputs, rather than on chain-of-thought tokens, making supervision easier to verify and less susceptible to reward hacking. CodeV represents visual tools as executable Python code, and TAPO assigns step-wise rewards based solely on the question and tool output, encouraging both necessary and evidence-consistent tool use. In a two-stage SFT+RL pipeline, CodeV achieves competitive or superior accuracy while substantially increasing faithful tool-use rates on related visual search benchmarks. Beyond visual search, CodeV attains strong performance on a range of multimodal reasoning and math benchmarks, suggesting that explicitly supervising intermediate tool behavior is crucial for building trustworthy, agentic visual reasoning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
Hou, Xinhai
Xu, Shaoyuan
Biyani, Manan
Li, Moyan
Liu, Jia
Hollon, Todd C.
Wang, Bryan
Computer Vision and Pattern Recognition
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful visual reasoning: models may invoke tools on irrelevant regions or ignore tool outputs entirely, yet still guess the correct answer. In this work, we first propose a faithfulness evaluation protocol that measures whether intermediate visual tool outputs (e.g., crops) actually contain the queried evidence. This reveals that recent visual agents achieve high final-answer accuracy but exhibit low rates of faithful tool-use on visual search benchmarks. We then introduce CodeV, a code-based visual agent trained with Tool-Aware Policy Optimization (TAPO). TAPO is a process-level RL framework that augments GRPO with dense rewards defined directly on visual tool inputs and outputs, rather than on chain-of-thought tokens, making supervision easier to verify and less susceptible to reward hacking. CodeV represents visual tools as executable Python code, and TAPO assigns step-wise rewards based solely on the question and tool output, encouraging both necessary and evidence-consistent tool use. In a two-stage SFT+RL pipeline, CodeV achieves competitive or superior accuracy while substantially increasing faithful tool-use rates on related visual search benchmarks. Beyond visual search, CodeV attains strong performance on a range of multimodal reasoning and math benchmarks, suggesting that explicitly supervising intermediate tool behavior is crucial for building trustworthy, agentic visual reasoning systems.
title CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19661