ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Tianyu, He, Chenwei, Hao, Xiangzhao, Wang, Tianyue, Guo, Jiarui, Guo, Haiyun, Qu, Leigang, Wang, Jinqiao, Chua, Tat-Seng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
di: Wang, Tianyue, et al.
Pubblicazione: (2026)
di: Wang, Tianyue, et al.
Pubblicazione: (2026)
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
di: Guo, Hongyu, et al.
Pubblicazione: (2025)
di: Guo, Hongyu, et al.
Pubblicazione: (2025)
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
di: Hao, Xiangzhao, et al.
Pubblicazione: (2026)
di: Hao, Xiangzhao, et al.
Pubblicazione: (2026)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
di: Chen, Yiyang, et al.
Pubblicazione: (2022)
di: Chen, Yiyang, et al.
Pubblicazione: (2022)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
di: Guo, Haiyun, et al.
Pubblicazione: (2025)
di: Guo, Haiyun, et al.
Pubblicazione: (2025)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
di: He, Chenwei, et al.
Pubblicazione: (2026)
di: He, Chenwei, et al.
Pubblicazione: (2026)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
di: Jin, Zhe, et al.
Pubblicazione: (2025)
di: Jin, Zhe, et al.
Pubblicazione: (2025)
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
di: Hao, Xiangzhao, et al.
Pubblicazione: (2025)
di: Hao, Xiangzhao, et al.
Pubblicazione: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
di: He, Jinghan, et al.
Pubblicazione: (2026)
di: He, Jinghan, et al.
Pubblicazione: (2026)
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
di: Wen, Haokun, et al.
Pubblicazione: (2024)
di: Wen, Haokun, et al.
Pubblicazione: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
di: He, Jinghan, et al.
Pubblicazione: (2024)
di: He, Jinghan, et al.
Pubblicazione: (2024)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
di: Hao, Xiangzhao, et al.
Pubblicazione: (2026)
di: Hao, Xiangzhao, et al.
Pubblicazione: (2026)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
di: Gao, Haowen, et al.
Pubblicazione: (2025)
di: Gao, Haowen, et al.
Pubblicazione: (2025)
VINCIE: Unlocking In-context Image Editing from Video
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
AAformer: Auto-Aligned Transformer for Person Re-Identification
di: Zhu, Kuan, et al.
Pubblicazione: (2021)
di: Zhu, Kuan, et al.
Pubblicazione: (2021)
Universal Scene Graph Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
di: Guo, Shasha, et al.
Pubblicazione: (2024)
di: Guo, Shasha, et al.
Pubblicazione: (2024)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)
A Taxation Perspective for Fair Re-ranking
di: Xu, Chen, et al.
Pubblicazione: (2024)
di: Xu, Chen, et al.
Pubblicazione: (2024)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
di: Zhang, An, et al.
Pubblicazione: (2024)
di: Zhang, An, et al.
Pubblicazione: (2024)
Distillation Enhanced Generative Retrieval
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
AUHead: Realistic Emotional Talking Head Generation via Action Units Control
di: Lyu, Jiayi, et al.
Pubblicazione: (2026)
di: Lyu, Jiayi, et al.
Pubblicazione: (2026)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
di: Fang, Xiang, et al.
Pubblicazione: (2026)
di: Fang, Xiang, et al.
Pubblicazione: (2026)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
di: Wang, Yabing, et al.
Pubblicazione: (2025)
di: Wang, Yabing, et al.
Pubblicazione: (2025)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
di: Ma, Weijian, et al.
Pubblicazione: (2026)
di: Ma, Weijian, et al.
Pubblicazione: (2026)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
XNLP: An Interactive Demonstration System for Universal Structured NLP
di: Fei, Hao, et al.
Pubblicazione: (2023)
di: Fei, Hao, et al.
Pubblicazione: (2023)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
di: Chu, Meng, et al.
Pubblicazione: (2025)
di: Chu, Meng, et al.
Pubblicazione: (2025)
Learning to Ask Critical Questions for Assisting Product Search
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
Rubric-based On-policy Distillation
di: Fang, Junfeng, et al.
Pubblicazione: (2026)
di: Fang, Junfeng, et al.
Pubblicazione: (2026)
FashionReGen: LLM-Empowered Fashion Report Generation
di: Ding, Yujuan, et al.
Pubblicazione: (2024)
di: Ding, Yujuan, et al.
Pubblicazione: (2024)
Exploring Training and Inference Scaling Laws in Generative Retrieval
di: Cai, Hongru, et al.
Pubblicazione: (2025)
di: Cai, Hongru, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
di: Wang, Tianyue, et al.
Pubblicazione: (2026) -
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
di: Guo, Hongyu, et al.
Pubblicazione: (2025) -
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
di: Hao, Xiangzhao, et al.
Pubblicazione: (2026) -
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
di: Chen, Yiyang, et al.
Pubblicazione: (2022) -
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
di: Li, Yongqi, et al.
Pubblicazione: (2024)