Benchmark Granularity and Model Robustness for Image-Text Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hendriksen, Mariya, Zhang, Shuo, Reinanda, Ridho, Yahya, Mohamed, Meij, Edgar, de Rijke, Maarten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Learned Sparse Retrieval with Probabilistic Expansion Control
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
von: Mezzi, Emanuele, et al.
Veröffentlicht: (2025)
von: Mezzi, Emanuele, et al.
Veröffentlicht: (2025)
Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
von: Yang, Wei, et al.
Veröffentlicht: (2025)
von: Yang, Wei, et al.
Veröffentlicht: (2025)
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
von: Long, Zijun, et al.
Veröffentlicht: (2025)
von: Long, Zijun, et al.
Veröffentlicht: (2025)
Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
von: Bleeker, Maurits, et al.
Veröffentlicht: (2024)
von: Bleeker, Maurits, et al.
Veröffentlicht: (2024)
CFIR: Fast and Effective Long-Text To Image Retrieval for Large Corpora
von: Long, Zijun, et al.
Veröffentlicht: (2024)
von: Long, Zijun, et al.
Veröffentlicht: (2024)
Content-Based Image Retrieval for Multi-Class Volumetric Radiology Images: A Benchmark Study
von: Jush, Farnaz Khun, et al.
Veröffentlicht: (2024)
von: Jush, Farnaz Khun, et al.
Veröffentlicht: (2024)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval
von: Deanda, Demetrio, et al.
Veröffentlicht: (2025)
von: Deanda, Demetrio, et al.
Veröffentlicht: (2025)
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
von: Sogi, Naoya, et al.
Veröffentlicht: (2025)
von: Sogi, Naoya, et al.
Veröffentlicht: (2025)
Are They the Same Picture? Adapting Concept Bottleneck Models for Human-AI Collaboration in Image Retrieval
von: Balloli, Vaibhav, et al.
Veröffentlicht: (2024)
von: Balloli, Vaibhav, et al.
Veröffentlicht: (2024)
Weakly Supervised Deep Hyperspherical Quantization for Image Retrieval
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
Retrieval-Guided Generation for Safer Histopathology Image Captioning
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
von: Hoq, Md. Enamul, et al.
Veröffentlicht: (2026)
Attributes Grouping and Mining Hashing for Fine-Grained Image Retrieval
von: Lu, Xin, et al.
Veröffentlicht: (2023)
von: Lu, Xin, et al.
Veröffentlicht: (2023)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
von: Jiang, Xintong, et al.
Veröffentlicht: (2024)
PHPQ: Pyramid Hybrid Pooling Quantization for Efficient Fine-Grained Image Retrieval
von: Zeng, Ziyun, et al.
Veröffentlicht: (2021)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2021)
Multi-task Cross-modal Learning for Chest X-ray Image Retrieval
von: Liang, Zhaohui, et al.
Veröffentlicht: (2026)
von: Liang, Zhaohui, et al.
Veröffentlicht: (2026)
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
von: Yang, Yuxin, et al.
Veröffentlicht: (2025)
von: Yang, Yuxin, et al.
Veröffentlicht: (2025)
Compressible and Searchable: AI-native Multi-Modal Retrieval System with Learned Image Compression
von: Luo, Jixiang
Veröffentlicht: (2024)
von: Luo, Jixiang
Veröffentlicht: (2024)
A Comprehensive Survey on Composed Image Retrieval
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
Content-based 3D Image Retrieval and a ColBERT-inspired Re-ranking for Tumor Flagging and Staging
von: Jush, Farnaz Khun, et al.
Veröffentlicht: (2025)
von: Jush, Farnaz Khun, et al.
Veröffentlicht: (2025)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024)
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
von: Guo, Yuanhe, et al.
Veröffentlicht: (2025)
von: Guo, Yuanhe, et al.
Veröffentlicht: (2025)
Image-Seeking Intent Prediction for Cross-Device Product Search
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2025)
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2025)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
von: Askari, Arian, et al.
Veröffentlicht: (2025)
von: Askari, Arian, et al.
Veröffentlicht: (2025)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024)
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
Self-supervised Graph Neural Network for Mechanical CAD Retrieval
von: Quan, Yuhan, et al.
Veröffentlicht: (2024)
von: Quan, Yuhan, et al.
Veröffentlicht: (2024)
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Open Multimodal Retrieval-Augmented Factual Image Generation
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2024)
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multimodal Learned Sparse Retrieval with Probabilistic Expansion Control
von: Nguyen, Thong, et al.
Veröffentlicht: (2024) -
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
von: Mezzi, Emanuele, et al.
Veröffentlicht: (2025) -
Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
von: Xu, Shicheng, et al.
Veröffentlicht: (2023) -
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
von: Yang, Wei, et al.
Veröffentlicht: (2025) -
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)