FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Williams-Lekuona, Mikel, Cosma, Georgina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
by: Williams-Lekuona, Mikel, et al.
Published: (2025)
by: Williams-Lekuona, Mikel, et al.
Published: (2025)
Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval
by: Liu, Delong, et al.
Published: (2023)
by: Liu, Delong, et al.
Published: (2023)
Direct content-based retrieval from music scores images
by: Luna-Barahona, Noelia, et al.
Published: (2026)
by: Luna-Barahona, Noelia, et al.
Published: (2026)
Acquisition of interpretable domain information during brain MR image harmonization for content-based image retrieval
by: Abe, Keima, et al.
Published: (2025)
by: Abe, Keima, et al.
Published: (2025)
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification
by: Mao, Chen, et al.
Published: (2024)
by: Mao, Chen, et al.
Published: (2024)
Domain-invariant feature learning in brain MR imaging for content-based image retrieval
by: Tobari, Shuya, et al.
Published: (2025)
by: Tobari, Shuya, et al.
Published: (2025)
MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark
by: Osmulski, Radek, et al.
Published: (2025)
by: Osmulski, Radek, et al.
Published: (2025)
Uncertainty-aware sign language video retrieval with probability distribution modeling
by: Wu, Xuan, et al.
Published: (2024)
by: Wu, Xuan, et al.
Published: (2024)
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
Introduction of a tree-based technique for efficient and real-time label retrieval in the object tracking system
by: Benrazek, Ala-Eddine, et al.
Published: (2022)
by: Benrazek, Ala-Eddine, et al.
Published: (2022)
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
by: Mezzi, Emanuele, et al.
Published: (2025)
by: Mezzi, Emanuele, et al.
Published: (2025)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
by: Luo, Enming, et al.
Published: (2024)
by: Luo, Enming, et al.
Published: (2024)
Improving fine-grained understanding in image-text pre-training
by: Bica, Ioana, et al.
Published: (2024)
by: Bica, Ioana, et al.
Published: (2024)
Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs
by: Zhang, Huaying, et al.
Published: (2024)
by: Zhang, Huaying, et al.
Published: (2024)
Chaining text-to-image and large language model: A novel approach for generating personalized e-commerce banners
by: Vashishtha, Shanu, et al.
Published: (2024)
by: Vashishtha, Shanu, et al.
Published: (2024)
CoLLM: A Large Language Model for Composed Image Retrieval
by: Huynh, Chuong, et al.
Published: (2025)
by: Huynh, Chuong, et al.
Published: (2025)
Benchmark Granularity and Model Robustness for Image-Text Retrieval
by: Hendriksen, Mariya, et al.
Published: (2024)
by: Hendriksen, Mariya, et al.
Published: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
by: Rosa, Kevin Dela
Published: (2025)
by: Rosa, Kevin Dela
Published: (2025)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
by: Lim, Ho Hung, et al.
Published: (2026)
by: Lim, Ho Hung, et al.
Published: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
by: Cui, Cheng, et al.
Published: (2026)
by: Cui, Cheng, et al.
Published: (2026)
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Weakly Supervised Deep Hyperspherical Quantization for Image Retrieval
by: Wang, Jinpeng, et al.
Published: (2024)
by: Wang, Jinpeng, et al.
Published: (2024)
Low-Data Classification of Historical Music Manuscripts: A Few-Shot Learning Approach
by: Shatri, Elona, et al.
Published: (2024)
by: Shatri, Elona, et al.
Published: (2024)
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
by: Horn, Pius, et al.
Published: (2025)
by: Horn, Pius, et al.
Published: (2025)
From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline
by: Shoilee, Sarah Binta Alam, et al.
Published: (2026)
by: Shoilee, Sarah Binta Alam, et al.
Published: (2026)
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
by: Rosa, Kevin Dela
Published: (2024)
by: Rosa, Kevin Dela
Published: (2024)
Accelerating Flood Warnings by 10 Hours: The Power of River Network Topology in AI-enhanced Flood Forecasting
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
by: Duan, Yicheng, et al.
Published: (2025)
by: Duan, Yicheng, et al.
Published: (2025)
Content-based 3D Image Retrieval and a ColBERT-inspired Re-ranking for Tumor Flagging and Staging
by: Jush, Farnaz Khun, et al.
Published: (2025)
by: Jush, Farnaz Khun, et al.
Published: (2025)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
by: Long, Zijun, et al.
Published: (2025)
by: Long, Zijun, et al.
Published: (2025)
Your Embedding Model is SMARTer Than You Think
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
PHPQ: Pyramid Hybrid Pooling Quantization for Efficient Fine-Grained Image Retrieval
by: Zeng, Ziyun, et al.
Published: (2021)
by: Zeng, Ziyun, et al.
Published: (2021)
Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Efficient Logic Gate Networks for Video Copy Detection
by: Fojcik, Katarzyna
Published: (2026)
by: Fojcik, Katarzyna
Published: (2026)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
by: Yin, Xinlei, et al.
Published: (2026)
by: Yin, Xinlei, et al.
Published: (2026)
Good Scores, Bad Data: A Metric for Multimodal Coherence
by: Srinivasan, Vasundra
Published: (2026)
by: Srinivasan, Vasundra
Published: (2026)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
by: Rao, Varun Nagaraj, et al.
Published: (2024)
by: Rao, Varun Nagaraj, et al.
Published: (2024)
Are They the Same Picture? Adapting Concept Bottleneck Models for Human-AI Collaboration in Image Retrieval
by: Balloli, Vaibhav, et al.
Published: (2024)
by: Balloli, Vaibhav, et al.
Published: (2024)
Similar Items
-
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
by: Williams-Lekuona, Mikel, et al.
Published: (2025) -
Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval
by: Liu, Delong, et al.
Published: (2023) -
Direct content-based retrieval from music scores images
by: Luna-Barahona, Noelia, et al.
Published: (2026) -
Acquisition of interpretable domain information during brain MR image harmonization for content-based image retrieval
by: Abe, Keima, et al.
Published: (2025) -
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
by: Zhang, Yao, et al.
Published: (2026)