Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Bao, Peijun, Kong, Chenqi, Shao, Zihao, Ng, Boon Poh, Er, Meng Hwa, Kot, Alex C. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimBase: A Simple Baseline for Temporal Video Grounding
by: Bao, Peijun, et al.
Published: (2024)
by: Bao, Peijun, et al.
Published: (2024)
From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
by: Bao, Peijun, et al.
Published: (2026)
by: Bao, Peijun, et al.
Published: (2026)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery Detection
by: Kong, Chenqi, et al.
Published: (2024)
by: Kong, Chenqi, et al.
Published: (2024)
Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method with Forgery Style Mixture
by: Kong, Chenqi, et al.
Published: (2024)
by: Kong, Chenqi, et al.
Published: (2024)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
by: Yang, Siyuan, et al.
Published: (2023)
by: Yang, Siyuan, et al.
Published: (2023)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
by: Liu, Weijia, et al.
Published: (2025)
by: Liu, Weijia, et al.
Published: (2025)
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior
by: Guo, Chen, et al.
Published: (2025)
by: Guo, Chen, et al.
Published: (2025)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking
by: Wu, Qiangqiang, et al.
Published: (2025)
by: Wu, Qiangqiang, et al.
Published: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
Towards Unstructured Unlabeled Optical Mocap: A Video Helps!
by: Milef, Nicholas, et al.
Published: (2024)
by: Milef, Nicholas, et al.
Published: (2024)
Object-Centric Framework for Video Moment Retrieval
by: Li, Zongyao, et al.
Published: (2025)
by: Li, Zongyao, et al.
Published: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
by: Luo, Dezhao, et al.
Published: (2024)
by: Luo, Dezhao, et al.
Published: (2024)
Fine-Tuning Video-Text Contrastive Model for Primate Behavior Retrieval from Unlabeled Raw Videos
by: Santo, Giulio Cesare Mastrocinque, et al.
Published: (2025)
by: Santo, Giulio Cesare Mastrocinque, et al.
Published: (2025)
Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering
by: Song, Peipei, et al.
Published: (2025)
by: Song, Peipei, et al.
Published: (2025)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack
by: Peng, Rongxuan, et al.
Published: (2025)
by: Peng, Rongxuan, et al.
Published: (2025)
HarmoVid: Relightful Video Portrait Harmonization
by: Choi, Jun Myeong, et al.
Published: (2026)
by: Choi, Jun Myeong, et al.
Published: (2026)
Automatic Retrieval of Specific Cows from Unlabeled Videos
by: Lyu, Jiawen, et al.
Published: (2025)
by: Lyu, Jiawen, et al.
Published: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025)
by: Ren, Zhongwei, et al.
Published: (2025)
S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing with Statistical Tokens
by: Cai, Rizhao, et al.
Published: (2023)
by: Cai, Rizhao, et al.
Published: (2023)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
by: Yoon, Sunjae, et al.
Published: (2022)
by: Yoon, Sunjae, et al.
Published: (2022)
Event-aware Video Corpus Moment Retrieval
by: Hou, Danyang, et al.
Published: (2024)
by: Hou, Danyang, et al.
Published: (2024)
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
by: Li, Zhaoxu, et al.
Published: (2026)
by: Li, Zhaoxu, et al.
Published: (2026)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
by: Liang, Feng, et al.
Published: (2023)
by: Liang, Feng, et al.
Published: (2023)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
Similar Items
-
SimBase: A Simple Baseline for Temporal Video Grounding
by: Bao, Peijun, et al.
Published: (2024) -
From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
by: Lu, Hui, et al.
Published: (2025) -
ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
by: Bao, Peijun, et al.
Published: (2026) -
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023) -
MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery Detection
by: Kong, Chenqi, et al.
Published: (2024)