Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Indrehus, Kjetil, Duric, Adrian, Choi, Changkyu, Ramezani-Kebrya, Ali
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914539384078336
author Indrehus, Kjetil
Duric, Adrian
Choi, Changkyu
Ramezani-Kebrya, Ali
author_facet Indrehus, Kjetil
Duric, Adrian
Choi, Changkyu
Ramezani-Kebrya, Ali
contents Document Visual Question Answering (DocVQA) requires vision-language models to reason not only about what information in a document is relevant to a question, but also where the answer is grounded on the page. Existing DocVQA models entangle question-relevant evidence and answer localization and operate largely as black boxes, offering limited means to verify how predictions depend on visual evidence. We propose CoExVQA, a self-explainable DocVQA framework with a grounded reasoning process through a chain-of-explanation design. CoExVQA first identifies question-relevant evidence, then explicitly localizes the answer region, and finally decodes the answer exclusively from the grounded region. Prediction via CoExVQA's chain-of-explanation enables direct inspection and verification of the reasoning process across modalities. Empirical results show that restricting decoding to grounded evidence achieves SotA explainable DocVQA performance on PFL-DocVQA, improving ANLS by 12% over the current explainable baselines while providing transparent and verifiable predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06058
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
Indrehus, Kjetil
Duric, Adrian
Choi, Changkyu
Ramezani-Kebrya, Ali
Machine Learning
Computer Vision and Pattern Recognition
Document Visual Question Answering (DocVQA) requires vision-language models to reason not only about what information in a document is relevant to a question, but also where the answer is grounded on the page. Existing DocVQA models entangle question-relevant evidence and answer localization and operate largely as black boxes, offering limited means to verify how predictions depend on visual evidence. We propose CoExVQA, a self-explainable DocVQA framework with a grounded reasoning process through a chain-of-explanation design. CoExVQA first identifies question-relevant evidence, then explicitly localizes the answer region, and finally decodes the answer exclusively from the grounded region. Prediction via CoExVQA's chain-of-explanation enables direct inspection and verification of the reasoning process across modalities. Empirical results show that restricting decoding to grounded evidence achieves SotA explainable DocVQA performance on PFL-DocVQA, improving ANLS by 12% over the current explainable baselines while providing transparent and verifiable predictions.
title Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.06058