Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Haoqiang, Wen, Haokun, Song, Xuemeng, Liu, Meng, Hu, Yupeng, Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comprehensive Survey on Composed Image Retrieval
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
Pseudo-triplet Guided Few-shot Composed Image Retrieval
von: Hou, Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Bohan, et al.
Veröffentlicht: (2024)
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
von: Chen, Zhiwei, et al.
Veröffentlicht: (2025)
von: Chen, Zhiwei, et al.
Veröffentlicht: (2025)
FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
Self-Training Boosted Multi-Factor Matching Network for Composed Image Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2023)
von: Wen, Haokun, et al.
Veröffentlicht: (2023)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
von: Duan, Yiqun, et al.
Veröffentlicht: (2025)
von: Duan, Yiqun, et al.
Veröffentlicht: (2025)
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2025)
von: Li, Zixu, et al.
Veröffentlicht: (2025)
OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval
von: Chen, Zhiwei, et al.
Veröffentlicht: (2025)
von: Chen, Zhiwei, et al.
Veröffentlicht: (2025)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2025)
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2025)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
Bringing Textual Prompt to AI-Generated Image Quality Assessment
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
Fine-grained Image Quality Assessment for Perceptual Image Restoration
von: Sheng, Xiangfei, et al.
Veröffentlicht: (2025)
von: Sheng, Xiangfei, et al.
Veröffentlicht: (2025)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozong, et al.
Veröffentlicht: (2024)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
von: Feng, X., et al.
Veröffentlicht: (2024)
von: Feng, X., et al.
Veröffentlicht: (2024)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
von: Zhang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Zhang, Wenfeng, et al.
Veröffentlicht: (2026)
TOL: Textual Localization with OpenStreetMap
von: Liao, Youqi, et al.
Veröffentlicht: (2026)
von: Liao, Youqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Comprehensive Survey on Composed Image Retrieval
von: Song, Xuemeng, et al.
Veröffentlicht: (2025) -
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026) -
Pseudo-triplet Guided Few-shot Composed Image Retrieval
von: Hou, Bohan, et al.
Veröffentlicht: (2024) -
HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
von: Chen, Zhiwei, et al.
Veröffentlicht: (2025) -
FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
von: Wen, Haokun, et al.
Veröffentlicht: (2026)