SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jain, Chelsi, Wu, Yiran, Zeng, Yifan, Liu, Jiale, Dai, S hengyu, Shao, Zhenwen, Wu, Qingyun, Wang, Huazheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913897942876160
author Jain, Chelsi
Wu, Yiran
Zeng, Yifan
Liu, Jiale
Dai, S hengyu
Shao, Zhenwen
Wu, Qingyun
Wang, Huazheng
author_facet Jain, Chelsi
Wu, Yiran
Zeng, Yifan
Liu, Jiale
Dai, S hengyu
Shao, Zhenwen
Wu, Qingyun
Wang, Huazheng
contents Document Visual Question Answering (DocVQA) is a practical yet challenging task, which is to ask questions based on documents while referring to multiple pages and different modalities of information, e.g, images and tables. To handle multi-modality, recent methods follow a similar Retrieval Augmented Generation (RAG) pipeline, but utilize Visual Language Models (VLMs) based embedding model to embed and retrieve relevant pages as images, and generate answers with VLMs that can accept an image as input. In this paper, we introduce SimpleDoc, a lightweight yet powerful retrieval - augmented framework for DocVQA. It boosts evidence page gathering by first retrieving candidates through embedding similarity and then filtering and re-ranking these candidates based on page summaries. A single VLM-based reasoner agent repeatedly invokes this dual-cue retriever, iteratively pulling fresh pages into a working memory until the question is confidently answered. SimpleDoc outperforms previous baselines by 3.2% on average on 4 DocVQA datasets with much fewer pages retrieved. Our code is available at https://github.com/ag2ai/SimpleDoc.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
Jain, Chelsi
Wu, Yiran
Zeng, Yifan
Liu, Jiale
Dai, S hengyu
Shao, Zhenwen
Wu, Qingyun
Wang, Huazheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Document Visual Question Answering (DocVQA) is a practical yet challenging task, which is to ask questions based on documents while referring to multiple pages and different modalities of information, e.g, images and tables. To handle multi-modality, recent methods follow a similar Retrieval Augmented Generation (RAG) pipeline, but utilize Visual Language Models (VLMs) based embedding model to embed and retrieve relevant pages as images, and generate answers with VLMs that can accept an image as input. In this paper, we introduce SimpleDoc, a lightweight yet powerful retrieval - augmented framework for DocVQA. It boosts evidence page gathering by first retrieving candidates through embedding similarity and then filtering and re-ranking these candidates based on page summaries. A single VLM-based reasoner agent repeatedly invokes this dual-cue retriever, iteratively pulling fresh pages into a working memory until the question is confidently answered. SimpleDoc outperforms previous baselines by 3.2% on average on 4 DocVQA datasets with much fewer pages retrieved. Our code is available at https://github.com/ag2ai/SimpleDoc.
title SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.14035