ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Shaofeng, Lei, Ting, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VQA support to Arabic Language Learning Educational Tool
by: Delassi, Khaled Bachir, et al.
Published: (2025)
by: Delassi, Khaled Bachir, et al.
Published: (2025)
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA
by: Thakrar, Karishma, et al.
Published: (2025)
by: Thakrar, Karishma, et al.
Published: (2025)
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025)
by: Madaka, Madhuri Latha, et al.
Published: (2025)
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA
by: Barezi, Elham J., et al.
Published: (2024)
by: Barezi, Elham J., et al.
Published: (2024)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
by: Jin, Zhenchao, et al.
Published: (2024)
by: Jin, Zhenchao, et al.
Published: (2024)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
by: Singh, Anshul, et al.
Published: (2025)
by: Singh, Anshul, et al.
Published: (2025)
GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
Kvasir-VQA: A Text-Image Pair GI Tract Dataset
by: Gautam, Sushant, et al.
Published: (2024)
by: Gautam, Sushant, et al.
Published: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025)
by: Niu, Lixing, et al.
Published: (2025)
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
by: Yang, Chen, et al.
Published: (2025)
by: Yang, Chen, et al.
Published: (2025)
Is ChatGPT-5 Ready for Mammogram VQA?
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
by: Chen, Yixiong, et al.
Published: (2026)
by: Chen, Yixiong, et al.
Published: (2026)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
by: Ma, Jiatong, et al.
Published: (2026)
by: Ma, Jiatong, et al.
Published: (2026)
Advancing Surgical VQA with Scene Graph Knowledge
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning
by: He, Yang, et al.
Published: (2026)
by: He, Yang, et al.
Published: (2026)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
by: Kang, Zeyi, et al.
Published: (2025)
by: Kang, Zeyi, et al.
Published: (2025)
MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction
by: Moradbeiki, Pardis, et al.
Published: (2025)
by: Moradbeiki, Pardis, et al.
Published: (2025)
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
by: Zhang, Situo, et al.
Published: (2026)
by: Zhang, Situo, et al.
Published: (2026)
Abduction of Domain Relationships from Data for VQA
by: Chowdhury, Al Mehdi Saadat, et al.
Published: (2025)
by: Chowdhury, Al Mehdi Saadat, et al.
Published: (2025)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
by: Ke, Yan, et al.
Published: (2025)
by: Ke, Yan, et al.
Published: (2025)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025)
by: Pandey, Ayush, et al.
Published: (2025)
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
by: Han, Hongyong, et al.
Published: (2025)
by: Han, Hongyong, et al.
Published: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
by: Tuong, Nguyen Anh, et al.
Published: (2026)
by: Tuong, Nguyen Anh, et al.
Published: (2026)
Brain-IT-VQA: From Brain Signals to Answers
by: Beliy, Roman, et al.
Published: (2026)
by: Beliy, Roman, et al.
Published: (2026)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images
by: Shen, Jialu, et al.
Published: (2026)
by: Shen, Jialu, et al.
Published: (2026)
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
by: Asadi, Mohammad, et al.
Published: (2026)
by: Asadi, Mohammad, et al.
Published: (2026)
MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
by: Gong, Siyu, et al.
Published: (2026)
by: Gong, Siyu, et al.
Published: (2026)
Similar Items
-
VQA support to Arabic Language Learning Educational Tool
by: Delassi, Khaled Bachir, et al.
Published: (2025) -
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA
by: Thakrar, Karishma, et al.
Published: (2025) -
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025) -
KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
by: Yang, Fan, et al.
Published: (2026) -
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA
by: Barezi, Elham J., et al.
Published: (2024)