Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Siting, Gao, Xiang, Du, Simon Shaolei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024)
by: Li, Siting, et al.
Published: (2024)
Embedding Geometries of Contrastive Language-Image Pre-Training
by: Chou, Jason Chuan-Chih, et al.
Published: (2024)
by: Chou, Jason Chuan-Chih, et al.
Published: (2024)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024)
by: Miranda, Imanol, et al.
Published: (2024)
Attribute Diversity Determines the Systematicity Gap in VQA
by: Berlot-Attwell, Ian, et al.
Published: (2023)
by: Berlot-Attwell, Ian, et al.
Published: (2023)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
by: Wu, Ning, et al.
Published: (2026)
by: Wu, Ning, et al.
Published: (2026)
MAGIC: Near-Optimal Data Attribution for Deep Learning
by: Ilyas, Andrew, et al.
Published: (2025)
by: Ilyas, Andrew, et al.
Published: (2025)
VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation
by: Rokuss, Maximilian, et al.
Published: (2025)
by: Rokuss, Maximilian, et al.
Published: (2025)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
by: Wu, Wentao, et al.
Published: (2023)
by: Wu, Wentao, et al.
Published: (2023)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
by: Meng, Debin, et al.
Published: (2025)
by: Meng, Debin, et al.
Published: (2025)
What Shape Is Optimal for Masks in Text Removal?
by: Nakada, Hyakka, et al.
Published: (2025)
by: Nakada, Hyakka, et al.
Published: (2025)
Efficient Domain Adaptation of Multimodal Embeddings using Constrastive Learning
by: Margaritis, Georgios, et al.
Published: (2025)
by: Margaritis, Georgios, et al.
Published: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Including Facial Expressions in Contextual Embeddings for Sign Language Generation
by: Viegas, Carla, et al.
Published: (2022)
by: Viegas, Carla, et al.
Published: (2022)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
by: Du, Yao, et al.
Published: (2026)
by: Du, Yao, et al.
Published: (2026)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
by: Xue, Yihao, et al.
Published: (2023)
by: Xue, Yihao, et al.
Published: (2023)
Reverse Stable Diffusion: What prompt was used to generate this image?
by: Croitoru, Florinel-Alin, et al.
Published: (2023)
by: Croitoru, Florinel-Alin, et al.
Published: (2023)
Hidden in Plain Sight -- Class Competition Focuses Attribution Maps
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
by: Jung, Hoin, et al.
Published: (2026)
by: Jung, Hoin, et al.
Published: (2026)
Teach CLIP to Develop a Number Sense for Ordinal Regression
by: Du, Yao, et al.
Published: (2024)
by: Du, Yao, et al.
Published: (2024)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
by: Yayavaram, Arnav, et al.
Published: (2025)
by: Yayavaram, Arnav, et al.
Published: (2025)
Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
by: Jin, Pengfei, et al.
Published: (2024)
by: Jin, Pengfei, et al.
Published: (2024)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2024)
by: Piergiovanni, AJ, et al.
Published: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
by: Lin, Sheng-Chieh, et al.
Published: (2024)
by: Lin, Sheng-Chieh, et al.
Published: (2024)
Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition
by: Tang, Minxue, et al.
Published: (2026)
by: Tang, Minxue, et al.
Published: (2026)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
by: Zhang, Letian, et al.
Published: (2023)
by: Zhang, Letian, et al.
Published: (2023)
CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
by: Vasilev, Viacheslav, et al.
Published: (2025)
by: Vasilev, Viacheslav, et al.
Published: (2025)
Optimizing CLIP Models for Image Retrieval with Maintained Joint-Embedding Alignment
by: Schall, Konstantin, et al.
Published: (2024)
by: Schall, Konstantin, et al.
Published: (2024)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
by: Zhu, Mengdan, et al.
Published: (2025)
by: Zhu, Mengdan, et al.
Published: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive Learning
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
Path Choice Matters for Clear Attribution in Path Methods
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2025)
by: Villegas, Danae Sánchez, et al.
Published: (2025)
Similar Items
-
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024) -
Embedding Geometries of Contrastive Language-Image Pre-Training
by: Chou, Jason Chuan-Chih, et al.
Published: (2024) -
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025) -
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
by: Miranda, Imanol, et al.
Published: (2024) -
Attribute Diversity Determines the Systematicity Gap in VQA
by: Berlot-Attwell, Ian, et al.
Published: (2023)