Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Zechen, Xiao, Tianjun, He, Tong, Wang, Pichao, Zhang, Zheng, Brox, Thomas, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
Factorized Visual Tokenization and Generation
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval
von: Sun, Xinyu, et al.
Veröffentlicht: (2026)
von: Sun, Xinyu, et al.
Veröffentlicht: (2026)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
von: Skow, Tyler, et al.
Veröffentlicht: (2026)
von: Skow, Tyler, et al.
Veröffentlicht: (2026)
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
von: Sun, Zengbao, et al.
Veröffentlicht: (2024)
von: Sun, Zengbao, et al.
Veröffentlicht: (2024)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
von: Long, Zijun, et al.
Veröffentlicht: (2025)
von: Long, Zijun, et al.
Veröffentlicht: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval
von: Zhang, Zhuocheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhuocheng, et al.
Veröffentlicht: (2026)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
von: Mounis, Mohamed Darwish, et al.
Veröffentlicht: (2026)
von: Mounis, Mohamed Darwish, et al.
Veröffentlicht: (2026)
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
von: Reddy, Arun, et al.
Veröffentlicht: (2025)
von: Reddy, Arun, et al.
Veröffentlicht: (2025)
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
von: Le, Hoang-Bao, et al.
Veröffentlicht: (2025)
von: Le, Hoang-Bao, et al.
Veröffentlicht: (2025)
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
von: Yang, Jiashu, et al.
Veröffentlicht: (2026)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
Composed Multi-modal Retrieval: A Survey of Approaches and Applications
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
Human-Oriented Image Retrieval System (HORSE): A Neuro-Symbolic Approach to Optimizing Retrieval of Previewed Images
von: Weinberg, Abraham Itzhak
Veröffentlicht: (2025)
von: Weinberg, Abraham Itzhak
Veröffentlicht: (2025)
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval
von: Byun, Jaeseok, et al.
Veröffentlicht: (2024)
von: Byun, Jaeseok, et al.
Veröffentlicht: (2024)
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2025)
ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification
von: Mao, Chen, et al.
Veröffentlicht: (2024)
von: Mao, Chen, et al.
Veröffentlicht: (2024)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
Offline Evaluation of Set-Based Text-to-Image Generation
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data
von: Lei, Tianhao, et al.
Veröffentlicht: (2026)
von: Lei, Tianhao, et al.
Veröffentlicht: (2026)
DEMO: A Statistical Perspective for Efficient Image-Text Matching
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
Benchmark Granularity and Model Robustness for Image-Text Retrieval
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
von: Li, Minghan, et al.
Veröffentlicht: (2026)
von: Li, Minghan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024) -
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024) -
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
von: Ning, Hailong, et al.
Veröffentlicht: (2025) -
Factorized Visual Tokenization and Generation
von: Bai, Zechen, et al.
Veröffentlicht: (2024) -
TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval
von: Sun, Xinyu, et al.
Veröffentlicht: (2026)