Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lompo, Boammani Aser, Haraoui, Marc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-objective Representation for Numbers in Clinical Narratives: A CamemBERT-Bio-Based Alternative to Large-Scale LLMs
von: Lompo, Boammani Aser, et al.
Veröffentlicht: (2024)
von: Lompo, Boammani Aser, et al.
Veröffentlicht: (2024)
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
von: Yang, Zheyuan, et al.
Veröffentlicht: (2026)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2026)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
Vero: An Open RL Recipe for General Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
CartoMapQA: A Fundamental Benchmark Dataset Evaluating Vision-Language Models on Cartographic Map Understanding
von: Ung, Huy Quang, et al.
Veröffentlicht: (2025)
von: Ung, Huy Quang, et al.
Veröffentlicht: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
von: Hwang, Taebaek, et al.
Veröffentlicht: (2025)
von: Hwang, Taebaek, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
von: Alonso, Iñigo, et al.
Veröffentlicht: (2025)
von: Alonso, Iñigo, et al.
Veröffentlicht: (2025)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
von: Huang, Yixu, et al.
Veröffentlicht: (2026)
von: Huang, Yixu, et al.
Veröffentlicht: (2026)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
von: Mayer, Julius, et al.
Veröffentlicht: (2025)
von: Mayer, Julius, et al.
Veröffentlicht: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
von: Kim, Taehee, et al.
Veröffentlicht: (2024)
von: Kim, Taehee, et al.
Veröffentlicht: (2024)
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
von: Pande, Nilay, et al.
Veröffentlicht: (2025)
von: Pande, Nilay, et al.
Veröffentlicht: (2025)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-objective Representation for Numbers in Clinical Narratives: A CamemBERT-Bio-Based Alternative to Large-Scale LLMs
von: Lompo, Boammani Aser, et al.
Veröffentlicht: (2024) -
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
von: Yang, Zheyuan, et al.
Veröffentlicht: (2026) -
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
von: Yu, Suhao, et al.
Veröffentlicht: (2025) -
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024) -
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)