Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Siyuan, Qu, Xiaoye, Li, Yafu, Zhu, Tong, He, Zefeng, Fu, Muxin, Liu, Daizong, Zheng, Wei-Long, Cheng, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913103444180992
author Huang, Siyuan
Qu, Xiaoye
Li, Yafu
Zhu, Tong
He, Zefeng
Fu, Muxin
Liu, Daizong
Zheng, Wei-Long
Cheng, Yu
author_facet Huang, Siyuan
Qu, Xiaoye
Li, Yafu
Zhu, Tong
He, Zefeng
Fu, Muxin
Liu, Daizong
Zheng, Wei-Long
Cheng, Yu
contents While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumulation of textual history expands the attention partition function, causing visual attention to decay inversely with generated sequence length. To counteract this, we propose Persistent Visual Memory (PVM), a lightweight learnable module designed to strengthen sustained, on-demand access to visual evidence. Integrated as a parallel branch alongside the Feed-Forward Network (FFN) in LVLMs, PVM establishes a distance-agnostic retrieval pathway that directly provides visual embeddings for enhanced visual perception, thereby structurally mitigating the signal suppression inherent to deep generation. Extensive experiments on Qwen3-VL models demonstrate that PVM brings notable improvements with negligible parameter overhead, delivering consistent average accuracy gains across both 4B and 8B scales, particularly in complex reasoning tasks that demand persistent visual perception. Furthermore, in-depth analysis reveals that PVM shows improved robustness in longer generations and accelerates internal prediction convergence.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00814
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
Huang, Siyuan
Qu, Xiaoye
Li, Yafu
Zhu, Tong
He, Zefeng
Fu, Muxin
Liu, Daizong
Zheng, Wei-Long
Cheng, Yu
Computer Vision and Pattern Recognition
Artificial Intelligence
While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumulation of textual history expands the attention partition function, causing visual attention to decay inversely with generated sequence length. To counteract this, we propose Persistent Visual Memory (PVM), a lightweight learnable module designed to strengthen sustained, on-demand access to visual evidence. Integrated as a parallel branch alongside the Feed-Forward Network (FFN) in LVLMs, PVM establishes a distance-agnostic retrieval pathway that directly provides visual embeddings for enhanced visual perception, thereby structurally mitigating the signal suppression inherent to deep generation. Extensive experiments on Qwen3-VL models demonstrate that PVM brings notable improvements with negligible parameter overhead, delivering consistent average accuracy gains across both 4B and 8B scales, particularly in complex reasoning tasks that demand persistent visual perception. Furthermore, in-depth analysis reveals that PVM shows improved robustness in longer generations and accelerates internal prediction convergence.
title Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.00814