Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Williams-Lekuona, Mikel, Cosma, Georgina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis
by: Williams-Lekuona, Mikel, et al.
Published: (2024)
by: Williams-Lekuona, Mikel, et al.
Published: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
by: Duan, Yue, et al.
Published: (2024)
by: Duan, Yue, et al.
Published: (2024)
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
by: Wang, Yuting, et al.
Published: (2023)
by: Wang, Yuting, et al.
Published: (2023)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
A Comprehensive Survey on Composed Image Retrieval
by: Song, Xuemeng, et al.
Published: (2025)
by: Song, Xuemeng, et al.
Published: (2025)
A Matter of Time: Revealing the Structure of Time in Vision-Language Models
by: Tekaya, Nidham, et al.
Published: (2025)
by: Tekaya, Nidham, et al.
Published: (2025)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
by: Zou, Qiang, et al.
Published: (2025)
by: Zou, Qiang, et al.
Published: (2025)
Infinite Video Understanding
by: Zhang, Dell, et al.
Published: (2025)
by: Zhang, Dell, et al.
Published: (2025)
Counteracting temporal attacks in Video Copy Detection
by: Fojcik, Katarzyna, et al.
Published: (2025)
by: Fojcik, Katarzyna, et al.
Published: (2025)
TelcoAI: Advancing 3GPP Technical Specification Search through Agentic Multi-Modal Retrieval-Augmented Generation
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
A Survey of Multimodal Composite Editing and Retrieval
by: Li, Suyan, et al.
Published: (2024)
by: Li, Suyan, et al.
Published: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
by: Gatti, Prajwal, et al.
Published: (2025)
by: Gatti, Prajwal, et al.
Published: (2025)
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
by: Jing, Xiaolun, et al.
Published: (2024)
by: Jing, Xiaolun, et al.
Published: (2024)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024)
by: Ju, Yeong-Joon, et al.
Published: (2024)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
Very Efficient Listwise Multimodal Reranking for Long Documents
by: Sun, Yiqun, et al.
Published: (2026)
by: Sun, Yiqun, et al.
Published: (2026)
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
by: Pegia, Maria-Eirini, et al.
Published: (2026)
by: Pegia, Maria-Eirini, et al.
Published: (2026)
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance
by: Thwal, Chu Myaet, et al.
Published: (2024)
by: Thwal, Chu Myaet, et al.
Published: (2024)
UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
by: Jiang, Haoyu, et al.
Published: (2024)
by: Jiang, Haoyu, et al.
Published: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
by: Yang, Jheng-Hong, et al.
Published: (2024)
by: Yang, Jheng-Hong, et al.
Published: (2024)
Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval
by: Moummad, Ilyass, et al.
Published: (2026)
by: Moummad, Ilyass, et al.
Published: (2026)
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
by: Yang, Mengzheng, et al.
Published: (2025)
by: Yang, Mengzheng, et al.
Published: (2025)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
by: Barrios, Wayner, et al.
Published: (2026)
by: Barrios, Wayner, et al.
Published: (2026)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
by: Messina, Nicola, et al.
Published: (2024)
by: Messina, Nicola, et al.
Published: (2024)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025)
by: Ning, Hailong, et al.
Published: (2025)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Introduction of a tree-based technique for efficient and real-time label retrieval in the object tracking system
by: Benrazek, Ala-Eddine, et al.
Published: (2022)
by: Benrazek, Ala-Eddine, et al.
Published: (2022)
LazyVLM: Neuro-Symbolic Approach to Video Analytics
by: Jian, Xiangru, et al.
Published: (2025)
by: Jian, Xiangru, et al.
Published: (2025)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
by: Zhang, Zhixin, et al.
Published: (2024)
by: Zhang, Zhixin, et al.
Published: (2024)
Similar Items
-
FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis
by: Williams-Lekuona, Mikel, et al.
Published: (2024) -
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
by: Duan, Yue, et al.
Published: (2024) -
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
by: Wang, Yuting, et al.
Published: (2023) -
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024) -
A Comprehensive Survey on Composed Image Retrieval
by: Song, Xuemeng, et al.
Published: (2025)