ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Yiming, Yang, Zhizheng, Liu, Zeming, Leng, Haitao, Liu, Shaoguo, Gao, Tingting, Liu, Qingjie, Wang, Yunhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912488966062080
author Lei, Yiming
Yang, Zhizheng
Liu, Zeming
Leng, Haitao
Liu, Shaoguo
Gao, Tingting
Liu, Qingjie
Wang, Yunhong
author_facet Lei, Yiming
Yang, Zhizheng
Liu, Zeming
Leng, Haitao
Liu, Shaoguo
Gao, Tingting
Liu, Qingjie
Wang, Yunhong
contents Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQFormer, which utilizes a memory block to enhance the presentation of contextual information. Furthermore, to facilitate further research, we carefully build a new multi-turn multi-modal dialogue dataset (TMDialog) for pre-training, instruction-tuning, and evaluation, which will be open-sourced lately. Compared with other multi-modal dialogue datasets, TMDialog contains longer conversations, which supports the research of multi-turn multi-modal dialogue. In addition, ContextQFormer is compared with three baselines on TMDialog and experimental results illustrate that ContextQFormer achieves an improvement of 2%-4% in available rate over baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23121
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
Lei, Yiming
Yang, Zhizheng
Liu, Zeming
Leng, Haitao
Liu, Shaoguo
Gao, Tingting
Liu, Qingjie
Wang, Yunhong
Computation and Language
Artificial Intelligence
Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQFormer, which utilizes a memory block to enhance the presentation of contextual information. Furthermore, to facilitate further research, we carefully build a new multi-turn multi-modal dialogue dataset (TMDialog) for pre-training, instruction-tuning, and evaluation, which will be open-sourced lately. Compared with other multi-modal dialogue datasets, TMDialog contains longer conversations, which supports the research of multi-turn multi-modal dialogue. In addition, ContextQFormer is compared with three baselines on TMDialog and experimental results illustrate that ContextQFormer achieves an improvement of 2%-4% in available rate over baselines.
title ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.23121