Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Yang, Liu, Yuqi, Jin, Qin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Efficient and High-Fidelity Omni Modality Retrieval
von: Huynh, Chuong, et al.
Veröffentlicht: (2026)
von: Huynh, Chuong, et al.
Veröffentlicht: (2026)
Towards Text-Image Interleaved Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
von: Kim, Seonok
Veröffentlicht: (2026)
von: Kim, Seonok
Veröffentlicht: (2026)
Benchmark Granularity and Model Robustness for Image-Text Retrieval
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
von: Hendriksen, Mariya, et al.
Veröffentlicht: (2024)
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
von: Sun, Zengbao, et al.
Veröffentlicht: (2024)
von: Sun, Zengbao, et al.
Veröffentlicht: (2024)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
von: Raja, Rahul, et al.
Veröffentlicht: (2025)
von: Raja, Rahul, et al.
Veröffentlicht: (2025)
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval
von: Deanda, Demetrio, et al.
Veröffentlicht: (2025)
von: Deanda, Demetrio, et al.
Veröffentlicht: (2025)
Multi-Vector Index Compression in Any Modality
von: Qin, Hanxiang, et al.
Veröffentlicht: (2026)
von: Qin, Hanxiang, et al.
Veröffentlicht: (2026)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
von: Patel, Piyushkumar
Veröffentlicht: (2025)
von: Patel, Piyushkumar
Veröffentlicht: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
von: Reddy, Arun, et al.
Veröffentlicht: (2025)
von: Reddy, Arun, et al.
Veröffentlicht: (2025)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
von: Xu, Zhengfei, et al.
Veröffentlicht: (2024)
von: Xu, Zhengfei, et al.
Veröffentlicht: (2024)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
Attention Grounded Enhancement for Visual Document Retrieval
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024) -
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024) -
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025) -
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024) -
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)