Towards General Continuous Memory for Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Wenyi, Song, Zixuan, Zhou, Kun, Shao, Yifei, Hu, Zhiting, Huang, Biwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912471024926720
author Wu, Wenyi
Song, Zixuan
Zhou, Kun
Shao, Yifei
Hu, Zhiting
Huang, Biwei
author_facet Wu, Wenyi
Song, Zixuan
Zhou, Kun
Shao, Yifei
Hu, Zhiting
Huang, Biwei
contents Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reasoning tasks that require multimodal or multilingual real-world knowledge. To support such capabilities, an external memory system that can efficiently provide relevant multimodal information is essential. Existing approaches generally concatenate image and text tokens into a long sequence as memory, which, however, may drastically increase context length and even degrade performance. In contrast, we propose using continuous memory, a compact set of dense embeddings to more effectively and efficiently represent multimodal and multilingual knowledge. Our key insight is that a VLM can serve as its own continuous memory encoder. We empirically show that this design improves performance on complex multimodal reasoning tasks. Building on this, we introduce a data-efficient and parameter-efficient method to fine-tune the VLM into a memory encoder, requiring only 1.2% of the model's parameters and a small corpus of 15.6K self-synthesized samples. Our approach CoMEM utilizes VLM's original capabilities to encode arbitrary multimodal and multilingual knowledge into just 8 continuous embeddings. Since the inference-time VLM remains frozen, our memory module is plug-and-play and can be flexibly integrated as needed. Extensive experiments across eight multimodal reasoning benchmarks demonstrate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards General Continuous Memory for Vision-Language Models
Wu, Wenyi
Song, Zixuan
Zhou, Kun
Shao, Yifei
Hu, Zhiting
Huang, Biwei
Machine Learning
Artificial Intelligence
Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reasoning tasks that require multimodal or multilingual real-world knowledge. To support such capabilities, an external memory system that can efficiently provide relevant multimodal information is essential. Existing approaches generally concatenate image and text tokens into a long sequence as memory, which, however, may drastically increase context length and even degrade performance. In contrast, we propose using continuous memory, a compact set of dense embeddings to more effectively and efficiently represent multimodal and multilingual knowledge. Our key insight is that a VLM can serve as its own continuous memory encoder. We empirically show that this design improves performance on complex multimodal reasoning tasks. Building on this, we introduce a data-efficient and parameter-efficient method to fine-tune the VLM into a memory encoder, requiring only 1.2% of the model's parameters and a small corpus of 15.6K self-synthesized samples. Our approach CoMEM utilizes VLM's original capabilities to encode arbitrary multimodal and multilingual knowledge into just 8 continuous embeddings. Since the inference-time VLM remains frozen, our memory module is plug-and-play and can be flexibly integrated as needed. Extensive experiments across eight multimodal reasoning benchmarks demonstrate the effectiveness of our approach.
title Towards General Continuous Memory for Vision-Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.17670