ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912488966062080 |
|---|---|
| author | Lei, Yiming Yang, Zhizheng Liu, Zeming Leng, Haitao Liu, Shaoguo Gao, Tingting Liu, Qingjie Wang, Yunhong |
| author_facet | Lei, Yiming Yang, Zhizheng Liu, Zeming Leng, Haitao Liu, Shaoguo Gao, Tingting Liu, Qingjie Wang, Yunhong |
| contents | Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQFormer, which utilizes a memory block to enhance the presentation of contextual information. Furthermore, to facilitate further research, we carefully build a new multi-turn multi-modal dialogue dataset (TMDialog) for pre-training, instruction-tuning, and evaluation, which will be open-sourced lately. Compared with other multi-modal dialogue datasets, TMDialog contains longer conversations, which supports the research of multi-turn multi-modal dialogue. In addition, ContextQFormer is compared with three baselines on TMDialog and experimental results illustrate that ContextQFormer achieves an improvement of 2%-4% in available rate over baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_23121 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Lei, Yiming Yang, Zhizheng Liu, Zeming Leng, Haitao Liu, Shaoguo Gao, Tingting Liu, Qingjie Wang, Yunhong Computation and Language Artificial Intelligence Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQFormer, which utilizes a memory block to enhance the presentation of contextual information. Furthermore, to facilitate further research, we carefully build a new multi-turn multi-modal dialogue dataset (TMDialog) for pre-training, instruction-tuning, and evaluation, which will be open-sourced lately. Compared with other multi-modal dialogue datasets, TMDialog contains longer conversations, which supports the research of multi-turn multi-modal dialogue. In addition, ContextQFormer is compared with three baselines on TMDialog and experimental results illustrate that ContextQFormer achieves an improvement of 2%-4% in available rate over baselines. |
| title | ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.23121 |