SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jin, Yiqiao, Kaur, Rachneet, Zeng, Zhen, Ganesh, Sumitra, Kumar, Srijan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908988534161408
author Jin, Yiqiao
Kaur, Rachneet
Zeng, Zhen
Ganesh, Sumitra
Kumar, Srijan
author_facet Jin, Yiqiao
Kaur, Rachneet
Zeng, Zhen
Ganesh, Sumitra
Kumar, Srijan
contents Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but it must balance limited effective context, redundant retrieved evidence, and the loss of fine-grained facts under aggressive compression. Pure compression-based approaches reduce input size but often discard fine-grained details essential for factual accuracy. We propose SARA, a hybrid RAG framework that targets answer quality under fixed token budgets by combining natural-language snippets with semantic compression vectors. SARA retains a small set of passages in text form to preserve entities and numerical values, compresses the remaining evidence into interpretable vectors for broader coverage, and uses those vectors for iterative evidence reranking. Across 9 datasets and 5 open-source LLMs spanning 3 model families (Mistral, Llama, and Gemma), SARA consistently improves answer relevance (+17.71), answer correctness (+13.72), and semantic similarity (+15.53), demonstrating the importance of integrating textual and compressed representations for robust, context-efficient RAG.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
Jin, Yiqiao
Kaur, Rachneet
Zeng, Zhen
Ganesh, Sumitra
Kumar, Srijan
Computation and Language
Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but it must balance limited effective context, redundant retrieved evidence, and the loss of fine-grained facts under aggressive compression. Pure compression-based approaches reduce input size but often discard fine-grained details essential for factual accuracy. We propose SARA, a hybrid RAG framework that targets answer quality under fixed token budgets by combining natural-language snippets with semantic compression vectors. SARA retains a small set of passages in text form to preserve entities and numerical values, compresses the remaining evidence into interpretable vectors for broader coverage, and uses those vectors for iterative evidence reranking. Across 9 datasets and 5 open-source LLMs spanning 3 model families (Mistral, Llama, and Gemma), SARA consistently improves answer relevance (+17.71), answer correctness (+13.72), and semantic similarity (+15.53), demonstrating the importance of integrating textual and compressed representations for robust, context-efficient RAG.
title SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
topic Computation and Language
url https://arxiv.org/abs/2510.26615