Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xianke, Liu, Daizong, Lou, Yushuo, Tan, Xin, Yang, Xun, Wang, Shuhui, Wang, Xun, Dong, Jianfeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910228941897728
author Chen, Xianke
Liu, Daizong
Lou, Yushuo
Tan, Xin
Yang, Xun
Wang, Shuhui
Wang, Xun
Dong, Jianfeng
author_facet Chen, Xianke
Liu, Daizong
Lou, Yushuo
Tan, Xin
Yang, Xun
Wang, Shuhui
Wang, Xun
Dong, Jianfeng
contents Different from traditional text-to-image retrieval tasks, chat-based image retrieval allows the human-interactive system to iteratively clarify and refine user intent through multi-round dialogue, thereby achieving more fine-grained retrieval results. The key challenge in this task lies in dynamically understanding and updating the user's query intent across dialogue rounds. Although existing works have achieved great performance on this new task, they simply handle history query information either by directly concatenating all previous queries into a long textual sequence or by relying on large language models to reconstruct the current query from history. Such strategies are computationally redundant and easily lead to inconsistent intent representations as the dialogue progresses. To alleviate these issues, this paper proposes a novel and efficient memory-based user intent updating framework for the chat-based image retrieval task, called Memory-Augmented Query Intent Understanding (MAQIU). It introduces a lightweight memorization module that dynamically aggregates and evolves the semantic representation of query intent across dialogues, while a memory recall mechanism is further employed to prevent intent forgetting and enhance long-term semantic integrity. In addition, MAQIU also integrates historical image retrieval results as visual guidance, allowing the model to strengthen cross-round correlations and refine current visual understanding. Extensive experiments demonstrate that MAQIU achieves substantial performance gains while maintaining high computational efficiency, reducing dialogue encoding FLOPs by 86.4\% compared with the prior baseline ChatIR. Source code is available at https://github.com/HuiGuanLab/MAQIU.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17365
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
Chen, Xianke
Liu, Daizong
Lou, Yushuo
Tan, Xin
Yang, Xun
Wang, Shuhui
Wang, Xun
Dong, Jianfeng
Computer Vision and Pattern Recognition
Different from traditional text-to-image retrieval tasks, chat-based image retrieval allows the human-interactive system to iteratively clarify and refine user intent through multi-round dialogue, thereby achieving more fine-grained retrieval results. The key challenge in this task lies in dynamically understanding and updating the user's query intent across dialogue rounds. Although existing works have achieved great performance on this new task, they simply handle history query information either by directly concatenating all previous queries into a long textual sequence or by relying on large language models to reconstruct the current query from history. Such strategies are computationally redundant and easily lead to inconsistent intent representations as the dialogue progresses. To alleviate these issues, this paper proposes a novel and efficient memory-based user intent updating framework for the chat-based image retrieval task, called Memory-Augmented Query Intent Understanding (MAQIU). It introduces a lightweight memorization module that dynamically aggregates and evolves the semantic representation of query intent across dialogues, while a memory recall mechanism is further employed to prevent intent forgetting and enhance long-term semantic integrity. In addition, MAQIU also integrates historical image retrieval results as visual guidance, allowing the model to strengthen cross-round correlations and refine current visual understanding. Extensive experiments demonstrate that MAQIU achieves substantial performance gains while maintaining high computational efficiency, reducing dialogue encoding FLOPs by 86.4\% compared with the prior baseline ChatIR. Source code is available at https://github.com/HuiGuanLab/MAQIU.
title Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17365