Where is this coming from? Making groundedness count in the evaluation of Document VQA models
Fuente:
arXiv
Saved in:
| Main Authors: | Nourbakhsh, Armineh, Parekh, Siddharth, Shetty, Pranav, Jin, Zhao, Shah, Sameena, Rose, Carolyn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TreeForm: End-to-end Annotation and Evaluation for Form Document Parsing
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
$R^2$-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
by: Wu, Zhen, et al.
Published: (2025)
by: Wu, Zhen, et al.
Published: (2025)
DocGraphLM: Documental Graph Language Model for Information Extraction
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025)
by: S, Santosh T. Y. S., et al.
Published: (2025)
Cognitive Agent Compilation for Explicit Problem Solver Modeling
by: Moon, Hyeongdon, et al.
Published: (2026)
by: Moon, Hyeongdon, et al.
Published: (2026)
CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis
by: Naik, Atharva, et al.
Published: (2026)
by: Naik, Atharva, et al.
Published: (2026)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
by: McCallum, Sabrina, et al.
Published: (2025)
by: McCallum, Sabrina, et al.
Published: (2025)
Advanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III
by: Shetty, Pranam, et al.
Published: (2025)
by: Shetty, Pranam, et al.
Published: (2025)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
by: Naik, Atharva, et al.
Published: (2024)
by: Naik, Atharva, et al.
Published: (2024)
BiCLIP: Domain Canonicalization via Structured Geometric Transformation
by: Mantini, Pranav, et al.
Published: (2026)
by: Mantini, Pranav, et al.
Published: (2026)
Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in Conversations
by: Dutt, Ritam, et al.
Published: (2024)
by: Dutt, Ritam, et al.
Published: (2024)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes
by: Wang, Rose E., et al.
Published: (2023)
by: Wang, Rose E., et al.
Published: (2023)
Where meaning lives: Layer-wise accessibility of psycholinguistic features in encoder and decoder language models
by: Tikhomirova, Taisiia, et al.
Published: (2026)
by: Tikhomirova, Taisiia, et al.
Published: (2026)
Supernova Event Dataset: Interpreting Large Language Models' Personality through Critical Event Analysis
by: Agarwal, Pranav, et al.
Published: (2025)
by: Agarwal, Pranav, et al.
Published: (2025)
AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection
by: Nourbakhsh, Aria, et al.
Published: (2026)
by: Nourbakhsh, Aria, et al.
Published: (2026)
Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation
by: Nourbakhsh, Aria, et al.
Published: (2026)
by: Nourbakhsh, Aria, et al.
Published: (2026)
Are LLMs good pragmatic speakers?
by: Jian, Mingyue, et al.
Published: (2024)
by: Jian, Mingyue, et al.
Published: (2024)
Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning
by: Betala, Siddharth, et al.
Published: (2024)
by: Betala, Siddharth, et al.
Published: (2024)
HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention
by: Goswami, Saumya, et al.
Published: (2025)
by: Goswami, Saumya, et al.
Published: (2025)
NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
by: Dutta, Aritra, et al.
Published: (2025)
by: Dutta, Aritra, et al.
Published: (2025)
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
by: Zhang, Qingru, et al.
Published: (2024)
by: Zhang, Qingru, et al.
Published: (2024)
Belief and Persuasion in Scientific Discourse on Social Media: A Study of the COVID-19 Pandemic
by: Alamir, Salwa, et al.
Published: (2024)
by: Alamir, Salwa, et al.
Published: (2024)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
by: Mehrafarin, Houman, et al.
Published: (2026)
by: Mehrafarin, Houman, et al.
Published: (2026)
Where is the Mind? Persona Vectors and LLM Individuation
by: Beckmann, Pierre, et al.
Published: (2026)
by: Beckmann, Pierre, et al.
Published: (2026)
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks
by: Parekh, Amit, et al.
Published: (2024)
by: Parekh, Amit, et al.
Published: (2024)
STRUX: An LLM for Decision-Making with Structured Explanations
by: Lu, Yiming, et al.
Published: (2024)
by: Lu, Yiming, et al.
Published: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Knowledge Generation for Zero-shot Knowledge-based VQA
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
by: Shah, Raj Sanjay, et al.
Published: (2025)
by: Shah, Raj Sanjay, et al.
Published: (2025)
Distilling Multi-Scale Knowledge for Event Temporal Relation Extraction
by: Yao, Hao-Ren, et al.
Published: (2022)
by: Yao, Hao-Ren, et al.
Published: (2022)
HALO: Hallucination Analysis and Learning Optimization to Empower LLMs with Retrieval-Augmented Context for Guided Clinical Decision Making
by: Anjum, Sumera, et al.
Published: (2024)
by: Anjum, Sumera, et al.
Published: (2024)
Learning to Trust the Crowd: A Multi-Model Consensus Reasoning Engine for Large Language Models
by: Kallem, Pranav
Published: (2026)
by: Kallem, Pranav
Published: (2026)
DIETA: A Decoder-only transformer-based model for Italian-English machine TrAnslation
by: Kasela, Pranav, et al.
Published: (2026)
by: Kasela, Pranav, et al.
Published: (2026)
Similar Items
-
TreeForm: End-to-end Annotation and Evaluation for Form Document Parsing
by: Zmigrod, Ran, et al.
Published: (2024) -
$R^2$-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
by: Wu, Zhen, et al.
Published: (2025) -
DocGraphLM: Documental Graph Language Model for Information Extraction
by: Wang, Dongsheng, et al.
Published: (2024) -
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024) -
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025)