Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Zongwan, Wen, Bingbing, Wang, Lucy Lu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
Estimating the Usefulness of Clarifying Questions and Answers for Conversational Search
by: Sekulić, Ivan, et al.
Published: (2024)
by: Sekulić, Ivan, et al.
Published: (2024)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
by: Baan, Joris, et al.
Published: (2026)
by: Baan, Joris, et al.
Published: (2026)
VAGUE: Visual Contexts Clarify Ambiguous Expressions
by: Nam, Heejeong, et al.
Published: (2024)
by: Nam, Heejeong, et al.
Published: (2024)
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
by: Xu, Chenjun, et al.
Published: (2026)
by: Xu, Chenjun, et al.
Published: (2026)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
Demystifying Reinforcement Learning in Agentic Reasoning
by: Yu, Zhaochen, et al.
Published: (2025)
by: Yu, Zhaochen, et al.
Published: (2025)
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
by: Feng, Andrew Zhuoer, et al.
Published: (2026)
by: Feng, Andrew Zhuoer, et al.
Published: (2026)
Self-Distilled Agentic Reinforcement Learning
by: Lu, Zhengxi, et al.
Published: (2026)
by: Lu, Zhengxi, et al.
Published: (2026)
Know Your Limits: A Survey of Abstention in Large Language Models
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
Agentic Reinforcement Learning with Implicit Step Rewards
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
Agentic Reinforcement Learning for Search is Unsafe
by: Yang, Yushi, et al.
Published: (2025)
by: Yang, Yushi, et al.
Published: (2025)
Context Embeddings for Efficient Answer Generation in RAG
by: Rau, David, et al.
Published: (2024)
by: Rau, David, et al.
Published: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
by: Zhu, Wenhui, et al.
Published: (2025)
by: Zhu, Wenhui, et al.
Published: (2025)
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
by: Takashiro, Shota, et al.
Published: (2024)
by: Takashiro, Shota, et al.
Published: (2024)
Answer Convergence as a Signal for Early Stopping in Reasoning
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
Journalism-Guided Agentic In-Context Learning for News Stance Detection
by: Lee, Dahyun, et al.
Published: (2025)
by: Lee, Dahyun, et al.
Published: (2025)
Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
by: Xin, Rihui, et al.
Published: (2025)
by: Xin, Rihui, et al.
Published: (2025)
Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning
by: Wang, Daoyu, et al.
Published: (2026)
by: Wang, Daoyu, et al.
Published: (2026)
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
by: Tan, Wentao, et al.
Published: (2025)
by: Tan, Wentao, et al.
Published: (2025)
From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning
by: zhang, Ranxu, et al.
Published: (2026)
by: zhang, Ranxu, et al.
Published: (2026)
Knowledge Generation for Zero-shot Knowledge-based VQA
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering
by: Liu, Yu, et al.
Published: (2026)
by: Liu, Yu, et al.
Published: (2026)
TR-ICRL: Test-Time Rethinking for In-Context Reinforcement Learning
by: Jiang, Wenxuan, et al.
Published: (2026)
by: Jiang, Wenxuan, et al.
Published: (2026)
Context Matters: Pushing the Boundaries of Open-Ended Answer Generation with Graph-Structured Knowledge Context
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Agentic Reinforcement Learning for Real-World Code Repair
by: Zhu, Siyu, et al.
Published: (2025)
by: Zhu, Siyu, et al.
Published: (2025)
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
On Many-Shot In-Context Learning for Long-Context Evaluation
by: Zou, Kaijian, et al.
Published: (2024)
by: Zou, Kaijian, et al.
Published: (2024)
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
by: Wang, Ziqing, et al.
Published: (2026)
by: Wang, Ziqing, et al.
Published: (2026)
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
by: Cheng, Xianfu, et al.
Published: (2025)
by: Cheng, Xianfu, et al.
Published: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
by: Gan, Yujian, et al.
Published: (2024)
by: Gan, Yujian, et al.
Published: (2024)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
by: Wen, Bingbing, et al.
Published: (2026)
by: Wen, Bingbing, et al.
Published: (2026)
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Similar Items
-
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
by: Wen, Bingbing, et al.
Published: (2024) -
Estimating the Usefulness of Clarifying Questions and Answers for Conversational Search
by: Sekulić, Ivan, et al.
Published: (2024) -
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
by: Baan, Joris, et al.
Published: (2026) -
VAGUE: Visual Contexts Clarify Ambiguous Expressions
by: Nam, Heejeong, et al.
Published: (2024) -
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
by: Xu, Chenjun, et al.
Published: (2026)