Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xianke, Liu, Daizong, Lou, Yushuo, Tan, Xin, Yang, Xun, Wang, Shuhui, Wang, Xun, Dong, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022)
by: Chen, Xianke, et al.
Published: (2022)
Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval
by: Dong, Jianfeng, et al.
Published: (2025)
by: Dong, Jianfeng, et al.
Published: (2025)
Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
by: Lin, Junan, et al.
Published: (2025)
by: Lin, Junan, et al.
Published: (2025)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
by: Cai, Rui, et al.
Published: (2024)
by: Cai, Rui, et al.
Published: (2024)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
by: Zhang, Long, et al.
Published: (2025)
by: Zhang, Long, et al.
Published: (2025)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
by: Song, Enxin, et al.
Published: (2023)
by: Song, Enxin, et al.
Published: (2023)
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO
by: Zhao, Fufangchen, et al.
Published: (2025)
by: Zhao, Fufangchen, et al.
Published: (2025)
Dual-stream Feature Augmentation for Domain Generalization
by: Wang, Shanshan, et al.
Published: (2024)
by: Wang, Shanshan, et al.
Published: (2024)
Representation Alignment Contrastive Regularization for Multi-Object Tracking
by: Liu, Zhonglin, et al.
Published: (2024)
by: Liu, Zhonglin, et al.
Published: (2024)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025)
by: Huang, Wencan, et al.
Published: (2025)
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering
by: Song, Peipei, et al.
Published: (2025)
by: Song, Peipei, et al.
Published: (2025)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
by: Cao, Yifei, et al.
Published: (2025)
by: Cao, Yifei, et al.
Published: (2025)
FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
by: Zhang, Rong, et al.
Published: (2025)
by: Zhang, Rong, et al.
Published: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
CLIP-based Camera-Agnostic Feature Learning for Intra-camera Person Re-Identification
by: Tan, Xuan, et al.
Published: (2024)
by: Tan, Xuan, et al.
Published: (2024)
Fine-Grained Controllable Apparel Showcase Image Generation via Garment-Centric Outpainting
by: Zhang, Rong, et al.
Published: (2025)
by: Zhang, Rong, et al.
Published: (2025)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
by: Xu, Zeyu, et al.
Published: (2025)
by: Xu, Zeyu, et al.
Published: (2025)
Efficient Visual Representation Learning with Heat Conduction Equation
by: Zhang, Zhemin, et al.
Published: (2024)
by: Zhang, Zhemin, et al.
Published: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Retrieval Augmented Image Harmonization
by: Wang, Haolin, et al.
Published: (2024)
by: Wang, Haolin, et al.
Published: (2024)
Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
by: Ji, Zhong, et al.
Published: (2023)
by: Ji, Zhong, et al.
Published: (2023)
NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries
by: Amemiya, Kanon, et al.
Published: (2026)
by: Amemiya, Kanon, et al.
Published: (2026)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network
by: Fang, Xiang, et al.
Published: (2024)
by: Fang, Xiang, et al.
Published: (2024)
Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
IntRec: Intent-based Retrieval with Contrastive Refinement
by: Shamsolmoali, Pourya, et al.
Published: (2026)
by: Shamsolmoali, Pourya, et al.
Published: (2026)
A Memory-Efficient Framework for Deformable Transformer with Neural Architecture Search
by: Mao, Wendong, et al.
Published: (2025)
by: Mao, Wendong, et al.
Published: (2025)
Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
by: Du, Ji, et al.
Published: (2025)
by: Du, Ji, et al.
Published: (2025)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2026)
by: Wu, Peiran, et al.
Published: (2026)
Retrieval Augmented Comic Image Generation
by: Shui, Yunhao, et al.
Published: (2025)
by: Shui, Yunhao, et al.
Published: (2025)
VideoChat: Chat-Centric Video Understanding
by: Li, KunChang, et al.
Published: (2023)
by: Li, KunChang, et al.
Published: (2023)
Data-Efficient Brushstroke Generation with Diffusion Models for Oil Painting
by: Qin, Dantong, et al.
Published: (2026)
by: Qin, Dantong, et al.
Published: (2026)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
by: Qi, Jingyuan, et al.
Published: (2025)
by: Qi, Jingyuan, et al.
Published: (2025)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Diffusion Models are Geometry Critics: Single Image 3D Editing Using Pre-Trained Diffusion Priors
by: Wang, Ruicheng, et al.
Published: (2024)
by: Wang, Ruicheng, et al.
Published: (2024)
Boosting Adversarial Transferability for Skeleton-based Action Recognition via Exploring the Model Posterior Space
by: Diao, Yunfeng, et al.
Published: (2024)
by: Diao, Yunfeng, et al.
Published: (2024)
Similar Items
-
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022) -
Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval
by: Dong, Jianfeng, et al.
Published: (2025) -
Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
by: Lin, Junan, et al.
Published: (2025) -
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
by: Cai, Rui, et al.
Published: (2024) -
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
by: Zhang, Long, et al.
Published: (2025)