MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yue, Hu, Jinwei, Lu, Qijia, Niu, Jiawei, Tan, Li, Yuan, Shuo, Yan, Ziyi, Jia, Yizhen, He, Qingzhi, Ge, Shiping, Chen, Ethan Q., Li, Wentong, Wang, Limin, Qin, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026)
by: Feng, Yue, et al.
Published: (2026)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025)
by: Peng, Liyang, et al.
Published: (2025)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
High-Dimensional Multi-Study Multi-Modality Covariate-Augmented Generalized Factor Model
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
by: Gan, Qijun, et al.
Published: (2024)
by: Gan, Qijun, et al.
Published: (2024)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024)
by: Tang, Yin, et al.
Published: (2024)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
by: Shuvo, Rezowan, et al.
Published: (2025)
by: Shuvo, Rezowan, et al.
Published: (2025)
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
by: Liu, Zhenghao, et al.
Published: (2025)
by: Liu, Zhenghao, et al.
Published: (2025)
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
by: Niu, Guanglin, et al.
Published: (2026)
by: Niu, Guanglin, et al.
Published: (2026)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
by: Li, Zizhong, et al.
Published: (2026)
by: Li, Zizhong, et al.
Published: (2026)
A Domain Decomposition Deep Neural Network Method with Multi-Activation Functions for Solving Elliptic and Parabolic Interface Problems
by: Zhai, Qijia
Published: (2025)
by: Zhai, Qijia
Published: (2025)
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
by: Hsu, Benjamin, et al.
Published: (2024)
by: Hsu, Benjamin, et al.
Published: (2024)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
High-Dimensional Covariate-Augmented Overdispersed Multi-Study Poisson Factor Model
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)
by: Helvaci, Halil Ismail, et al.
Published: (2024)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
by: Li, Qian, et al.
Published: (2024)
by: Li, Qian, et al.
Published: (2024)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
Prognostic Value of Combined Detection of MMP‐7 and ALP Levels in Children With Biliary Atresia Post‐Kasai Surgery
by: Qingzhi Li, et al.
Published: (2025)
by: Qingzhi Li, et al.
Published: (2025)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
MultiTEND: A Multilingual Benchmark for Natural Language to NoSQL Query Translation
by: Qin, Zhiqian, et al.
Published: (2025)
by: Qin, Zhiqian, et al.
Published: (2025)
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
by: Du, Yang, et al.
Published: (2024)
by: Du, Yang, et al.
Published: (2024)
Cross-Modal Navigation with Multi-Agent Reinforcement Learning
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
Multi‐Mode/Signal Biosensors: Electrochemical Integrated Sensing Techniques
by: Qingzhi Han, et al.
Published: (2024)
by: Qingzhi Han, et al.
Published: (2024)
MultiModal Action Conditioned Video Generation
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
by: Wang, Lintao, et al.
Published: (2025)
by: Wang, Lintao, et al.
Published: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
RAG over Tables: Hierarchical Memory Index, Multi-Stage Retrieval, and Benchmarking
by: Zou, Jiaru, et al.
Published: (2025)
by: Zou, Jiaru, et al.
Published: (2025)
MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module Plugin
by: Zhou, Tianshuo, et al.
Published: (2023)
by: Zhou, Tianshuo, et al.
Published: (2023)
Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
by: Yu, Jiashuo, et al.
Published: (2025)
by: Yu, Jiashuo, et al.
Published: (2025)
Metachromatic Butterfly Bile Pigments for Multi‐Level Optical Security Films
by: Limin Wang, et al.
Published: (2026)
by: Limin Wang, et al.
Published: (2026)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
by: Zhang, Jiaming, et al.
Published: (2023)
by: Zhang, Jiaming, et al.
Published: (2023)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
by: Kandhare, Mahesh, et al.
Published: (2024)
by: Kandhare, Mahesh, et al.
Published: (2024)
Similar Items
-
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026) -
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
by: Peng, Liyang, et al.
Published: (2025) -
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023) -
High-Dimensional Multi-Study Multi-Modality Covariate-Augmented Generalized Factor Model
by: Liu, Wei, et al.
Published: (2025) -
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
by: Gan, Qijun, et al.
Published: (2024)