Saved in:
Bibliographic Details
Main Authors: Yang, Tianyu, Ruas, Terry, Tian, Yijun, Wahle, Jan Philip, Kurzawe, Daniel, Gipp, Bela
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.25668
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914122067607552
author Yang, Tianyu
Ruas, Terry
Tian, Yijun
Wahle, Jan Philip
Kurzawe, Daniel
Gipp, Bela
author_facet Yang, Tianyu
Ruas, Terry
Tian, Yijun
Wahle, Jan Philip
Kurzawe, Daniel
Gipp, Bela
contents Vision-language models (VLMs) excel at interpreting text-rich images but struggle with long, visually complex documents that demand analysis and integration of information spread across multiple pages. Existing approaches typically rely on fixed reasoning templates or rigid pipelines, which force VLMs into a passive role and hinder both efficiency and generalization. We present Active Long-DocumEnt Navigation (ALDEN), a multi-turn reinforcement learning framework that fine-tunes VLMs as interactive agents capable of actively navigating long, visually rich documents. ALDEN introduces a novel fetch action that directly accesses the page by index, complementing the classic search action and better exploiting document structure. For dense process supervision and efficient training, we propose a rule-based cross-level reward that provides both turn- and token-level signals. To address the empirically observed training instability caused by numerous visual tokens from long documents, we further propose a visual-semantic anchoring mechanism that applies a dual-path KL-divergence constraint to stabilize visual and textual representations separately during training. Trained on a corpus constructed from three open-source datasets, ALDEN achieves state-of-the-art performance on five long-document benchmarks. Overall, ALDEN marks a step beyond passive document reading toward agents that autonomously navigate and reason across long, visually rich documents, offering a robust path to more accurate and efficient long-document understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25668
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
Yang, Tianyu
Ruas, Terry
Tian, Yijun
Wahle, Jan Philip
Kurzawe, Daniel
Gipp, Bela
Artificial Intelligence
Multimedia
Vision-language models (VLMs) excel at interpreting text-rich images but struggle with long, visually complex documents that demand analysis and integration of information spread across multiple pages. Existing approaches typically rely on fixed reasoning templates or rigid pipelines, which force VLMs into a passive role and hinder both efficiency and generalization. We present Active Long-DocumEnt Navigation (ALDEN), a multi-turn reinforcement learning framework that fine-tunes VLMs as interactive agents capable of actively navigating long, visually rich documents. ALDEN introduces a novel fetch action that directly accesses the page by index, complementing the classic search action and better exploiting document structure. For dense process supervision and efficient training, we propose a rule-based cross-level reward that provides both turn- and token-level signals. To address the empirically observed training instability caused by numerous visual tokens from long documents, we further propose a visual-semantic anchoring mechanism that applies a dual-path KL-divergence constraint to stabilize visual and textual representations separately during training. Trained on a corpus constructed from three open-source datasets, ALDEN achieves state-of-the-art performance on five long-document benchmarks. Overall, ALDEN marks a step beyond passive document reading toward agents that autonomously navigate and reason across long, visually rich documents, offering a robust path to more accurate and efficient long-document understanding.
title ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
topic Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2510.25668