Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Kawasaki, Haruka, Tanaka, Ryota, Nishida, Kyosuke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
por: Tanaka, Ryota, et al.
Publicado: (2024)
por: Tanaka, Ryota, et al.
Publicado: (2024)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
por: Tanaka, Ryota, et al.
Publicado: (2025)
por: Tanaka, Ryota, et al.
Publicado: (2025)
Internalized Reasoning for Long-Context Visual Document Understanding
por: Veselka, Austin
Publicado: (2026)
por: Veselka, Austin
Publicado: (2026)
Long Story Short: Story-level Video Understanding from 20K Short Films
por: Ghermi, Ridouane, et al.
Publicado: (2024)
por: Ghermi, Ridouane, et al.
Publicado: (2024)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
por: Vaishnav, Mohit, et al.
Publicado: (2026)
por: Vaishnav, Mohit, et al.
Publicado: (2026)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
por: Han, Jiaming, et al.
Publicado: (2025)
por: Han, Jiaming, et al.
Publicado: (2025)
Understanding Figurative Meaning through Explainable Visual Entailment
por: Saakyan, Arkadiy, et al.
Publicado: (2024)
por: Saakyan, Arkadiy, et al.
Publicado: (2024)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
por: Tartaglini, Alexa R., et al.
Publicado: (2025)
por: Tartaglini, Alexa R., et al.
Publicado: (2025)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
por: Seoh, Ronald, et al.
Publicado: (2025)
por: Seoh, Ronald, et al.
Publicado: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
por: Wang, Dianyi, et al.
Publicado: (2025)
por: Wang, Dianyi, et al.
Publicado: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
por: Wu, Chengyue, et al.
Publicado: (2024)
por: Wu, Chengyue, et al.
Publicado: (2024)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
por: Zhou, Kaiwen, et al.
Publicado: (2023)
por: Zhou, Kaiwen, et al.
Publicado: (2023)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
por: Fujitake, Masato
Publicado: (2024)
por: Fujitake, Masato
Publicado: (2024)
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
por: Jang, Jihyoung, et al.
Publicado: (2026)
por: Jang, Jihyoung, et al.
Publicado: (2026)
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
por: Zhang, Xianqi
Publicado: (2026)
por: Zhang, Xianqi
Publicado: (2026)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
por: Ni, Feng, et al.
Publicado: (2025)
por: Ni, Feng, et al.
Publicado: (2025)
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
por: Son, Jaemin, et al.
Publicado: (2025)
por: Son, Jaemin, et al.
Publicado: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
por: Huang, Kui, et al.
Publicado: (2025)
por: Huang, Kui, et al.
Publicado: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
por: Lin, Bin, et al.
Publicado: (2025)
por: Lin, Bin, et al.
Publicado: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
por: Huang, Kung-Hsiang, et al.
Publicado: (2025)
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
por: Zhou, Yuqi, et al.
Publicado: (2025)
por: Zhou, Yuqi, et al.
Publicado: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
por: Atuhurra, Jesse, et al.
Publicado: (2025)
por: Atuhurra, Jesse, et al.
Publicado: (2025)
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
por: Zhao, Yiming, et al.
Publicado: (2025)
por: Zhao, Yiming, et al.
Publicado: (2025)
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
por: Lan, Jian, et al.
Publicado: (2024)
por: Lan, Jian, et al.
Publicado: (2024)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
por: Li, Yuan, et al.
Publicado: (2026)
por: Li, Yuan, et al.
Publicado: (2026)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
por: Wang, Qiuchen, et al.
Publicado: (2025)
por: Wang, Qiuchen, et al.
Publicado: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
por: Zhu, Zichen, et al.
Publicado: (2024)
por: Zhu, Zichen, et al.
Publicado: (2024)
Benchmarking Vision Language Models for Cultural Understanding
por: Nayak, Shravan, et al.
Publicado: (2024)
por: Nayak, Shravan, et al.
Publicado: (2024)
Adaptive Greedy Frame Selection for Long Video Understanding
por: Huang, Yuning, et al.
Publicado: (2026)
por: Huang, Yuning, et al.
Publicado: (2026)
MLVU: Benchmarking Multi-task Long Video Understanding
por: Zhou, Junjie, et al.
Publicado: (2024)
por: Zhou, Junjie, et al.
Publicado: (2024)
Can Vision Language Models Understand Mimed Actions?
por: Cho, Hyundong, et al.
Publicado: (2025)
por: Cho, Hyundong, et al.
Publicado: (2025)
EgoNormia: Benchmarking Physical Social Norm Understanding
por: Rezaei, MohammadHossein, et al.
Publicado: (2025)
por: Rezaei, MohammadHossein, et al.
Publicado: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
por: Yilmaz, Nilay, et al.
Publicado: (2025)
por: Yilmaz, Nilay, et al.
Publicado: (2025)
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
por: Huybrechts, Goeric, et al.
Publicado: (2025)
por: Huybrechts, Goeric, et al.
Publicado: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
por: Takishita, Sho, et al.
Publicado: (2025)
por: Takishita, Sho, et al.
Publicado: (2025)
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
por: Lee, Jewon, et al.
Publicado: (2025)
por: Lee, Jewon, et al.
Publicado: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
por: Saxena, Rohit, et al.
Publicado: (2025)
por: Saxena, Rohit, et al.
Publicado: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
Ejemplares similares
-
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
por: Tanaka, Ryota, et al.
Publicado: (2024) -
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
por: Tanaka, Ryota, et al.
Publicado: (2025) -
Internalized Reasoning for Long-Context Visual Document Understanding
por: Veselka, Austin
Publicado: (2026) -
Long Story Short: Story-level Video Understanding from 20K Short Films
por: Ghermi, Ridouane, et al.
Publicado: (2024) -
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
por: Vaishnav, Mohit, et al.
Publicado: (2026)