Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Pegia, Maria-Eirini, Stefanopoulos, Dimitrios, Jónsson, Björn Þór, Moumtzidou, Anastasia, Gialampoukidis, Ilias, Vrochidis, Stefanos, Kompatsiaris, Ioannis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
by: Papadimitriou, Ioannis, et al.
Published: (2024)
by: Papadimitriou, Ioannis, et al.
Published: (2024)
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
by: Tran, Allie, et al.
Published: (2025)
by: Tran, Allie, et al.
Published: (2025)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
by: Wu, Jiaxin, et al.
Published: (2025)
by: Wu, Jiaxin, et al.
Published: (2025)
VCR: Video representation for Contextual Retrieval
by: Nir, Oron, et al.
Published: (2024)
by: Nir, Oron, et al.
Published: (2024)
Assessment of Oil Spill Dispersion and Weathering Processes in Saronic Gulf
by: Papaioannou, Vassilios, et al.
Published: (2025)
by: Papaioannou, Vassilios, et al.
Published: (2025)
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
by: Fu, Junchen, et al.
Published: (2026)
by: Fu, Junchen, et al.
Published: (2026)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
by: Jing, Xiaolun, et al.
Published: (2024)
by: Jing, Xiaolun, et al.
Published: (2024)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2024)
by: Xiao, Jian, et al.
Published: (2024)
InfoCIR: Multimedia Analysis for Composed Image Retrieval
by: Dravilas, Ioannis, et al.
Published: (2026)
by: Dravilas, Ioannis, et al.
Published: (2026)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2025)
by: Xiao, Jian, et al.
Published: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Performance Evaluation in Multimedia Retrieval
by: Sauter, Loris, et al.
Published: (2024)
by: Sauter, Loris, et al.
Published: (2024)
Multimodal Learned Sparse Retrieval for Image Suggestion
by: Nguyen, Thong, et al.
Published: (2024)
by: Nguyen, Thong, et al.
Published: (2024)
The Curious Case of High-Dimensional Indexing as a File Structure: A Case Study of eCP-FS
by: Khan, Omar Shahbaz, et al.
Published: (2025)
by: Khan, Omar Shahbaz, et al.
Published: (2025)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023)
by: Wang, Tianshi, et al.
Published: (2023)
Enhancing Image-Text Matching with Adaptive Feature Aggregation
by: Wang, Zuhui, et al.
Published: (2024)
by: Wang, Zuhui, et al.
Published: (2024)
Results of the 2025 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
Results of the 2024 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2024)
by: Rossetto, Luca, et al.
Published: (2024)
MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Balancing Semantic Relevance and Engagement in Related Video Recommendations
by: Jaspal, Amit, et al.
Published: (2025)
by: Jaspal, Amit, et al.
Published: (2025)
U-Sticker: A Large-Scale Multi-Domain User Sticker Dataset for Retrieval and Personalization
by: Chee, Heng Er Metilda, et al.
Published: (2025)
by: Chee, Heng Er Metilda, et al.
Published: (2025)
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025)
by: Ma, Zhixin, et al.
Published: (2025)
Leveraging User-Generated Metadata of Online Videos for Cover Song Identification
by: Hachmeier, Simon, et al.
Published: (2024)
by: Hachmeier, Simon, et al.
Published: (2024)
Small Stickers, Big Meanings: A Multilingual Sticker Semantic Understanding Dataset with a Gamified Approach
by: Chee, Heng Er Metilda, et al.
Published: (2025)
by: Chee, Heng Er Metilda, et al.
Published: (2025)
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
by: Sun, Huatuan, et al.
Published: (2025)
by: Sun, Huatuan, et al.
Published: (2025)
Interactive Multi-Turn Retrieval for Health Videos
by: Wu, Chengzheng, et al.
Published: (2026)
by: Wu, Chengzheng, et al.
Published: (2026)
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
by: Ning, Hailong, et al.
Published: (2025)
by: Ning, Hailong, et al.
Published: (2025)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
by: Messina, Nicola, et al.
Published: (2024)
by: Messina, Nicola, et al.
Published: (2024)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
VKIE: The Application of Key Information Extraction on Video Text
by: An, Siyu, et al.
Published: (2023)
by: An, Siyu, et al.
Published: (2023)
Music4All A+A: A Multimodal Dataset for Music Information Retrieval Tasks
by: Geiger, Jonas, et al.
Published: (2025)
by: Geiger, Jonas, et al.
Published: (2025)
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
On the Brittleness of CLIP Text Encoders
by: Tran, Allie, et al.
Published: (2025)
by: Tran, Allie, et al.
Published: (2025)
A Comprehensive Survey on Composed Image Retrieval
by: Song, Xuemeng, et al.
Published: (2025)
by: Song, Xuemeng, et al.
Published: (2025)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
Multimodal Information Retrieval for Open World with Edit Distance Weak Supervision
by: Solaiman, KMA, et al.
Published: (2025)
by: Solaiman, KMA, et al.
Published: (2025)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
Similar Items
-
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
by: Papadimitriou, Ioannis, et al.
Published: (2024) -
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
by: Tran, Allie, et al.
Published: (2025) -
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
by: Wu, Jiaxin, et al.
Published: (2025) -
VCR: Video representation for Contextual Retrieval
by: Nir, Oron, et al.
Published: (2024) -
Assessment of Oil Spill Dispersion and Weathering Processes in Saronic Gulf
by: Papaioannou, Vassilios, et al.
Published: (2025)