Object-Centric Framework for Video Moment Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zongyao, Wong, Yongkang, Yamazaki, Satoshi, Liu, Jianquan, Kankanhalli, Mohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
by: Li, Zongyao, et al.
Published: (2025)
by: Li, Zongyao, et al.
Published: (2025)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Learning to Predict Gradients for Semi-Supervised Continual Learning
by: Luo, Yan, et al.
Published: (2022)
by: Luo, Yan, et al.
Published: (2022)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
by: Chai, Zenghao, et al.
Published: (2024)
by: Chai, Zenghao, et al.
Published: (2024)
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
by: Guo, Yangyang, et al.
Published: (2023)
by: Guo, Yangyang, et al.
Published: (2023)
MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
by: Chai, Zenghao, et al.
Published: (2025)
by: Chai, Zenghao, et al.
Published: (2025)
SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency
by: Guo, Yangyang, et al.
Published: (2024)
by: Guo, Yangyang, et al.
Published: (2024)
MCM: Multi-condition Motion Synthesis Framework
by: Ling, Zeyu, et al.
Published: (2024)
by: Ling, Zeyu, et al.
Published: (2024)
MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Bridging the Intent Gap: Knowledge-Enhanced Visual Generation
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
Finetuning Text-to-Image Diffusion Models for Fairness
by: Shen, Xudong, et al.
Published: (2023)
by: Shen, Xudong, et al.
Published: (2023)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
by: Liu, Weijia, et al.
Published: (2025)
by: Liu, Weijia, et al.
Published: (2025)
MCM: Multi-condition Motion Synthesis Framework for Multi-scenario
by: Ling, Zeyu, et al.
Published: (2023)
by: Ling, Zeyu, et al.
Published: (2023)
Joint Vision-Language Social Bias Removal for CLIP
by: Zhang, Haoyu, et al.
Published: (2024)
by: Zhang, Haoyu, et al.
Published: (2024)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
Aggregating Diverse Cue Experts for AI-Generated Image Detection
by: Tan, Lei, et al.
Published: (2026)
by: Tan, Lei, et al.
Published: (2026)
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
by: Choong, Wey Yeh, et al.
Published: (2024)
by: Choong, Wey Yeh, et al.
Published: (2024)
Cycle Consistency in Video Object-Centric Learning
by: Zhao, Rongzhen, et al.
Published: (2026)
by: Zhao, Rongzhen, et al.
Published: (2026)
Video Spatial Reasoning with Object-Centric 3D Rollout
by: Tang, Haoran, et al.
Published: (2025)
by: Tang, Haoran, et al.
Published: (2025)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
by: Luo, Dezhao, et al.
Published: (2024)
by: Luo, Dezhao, et al.
Published: (2024)
FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
by: Wang, Tianyi, et al.
Published: (2025)
by: Wang, Tianyi, et al.
Published: (2025)
GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
by: Jeon, Mingyu, et al.
Published: (2026)
by: Jeon, Mingyu, et al.
Published: (2026)
SLVideo: A Sign Language Video Moment Retrieval Framework
by: Martins, Gonçalo Vinagre, et al.
Published: (2024)
by: Martins, Gonçalo Vinagre, et al.
Published: (2024)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)
by: Liu, Weijia, et al.
Published: (2024)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
Object-Centric Open-Vocabulary Image-Retrieval with Aggregated Features
by: Levi, Hila, et al.
Published: (2023)
by: Levi, Hila, et al.
Published: (2023)
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
by: Ding, Yiming, et al.
Published: (2026)
by: Ding, Yiming, et al.
Published: (2026)
Compositional Video Synthesis by Temporal Object-Centric Learning
by: Akan, Adil Kaan, et al.
Published: (2025)
by: Akan, Adil Kaan, et al.
Published: (2025)
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
by: Bao, Peijun, et al.
Published: (2024)
by: Bao, Peijun, et al.
Published: (2024)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
by: Yoon, Sunjae, et al.
Published: (2022)
by: Yoon, Sunjae, et al.
Published: (2022)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
by: Zhao, Henghao, et al.
Published: (2023)
by: Zhao, Henghao, et al.
Published: (2023)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Efficient Oriented Object Detection with Enhanced Small Object Recognition in Aerial Images
by: Shi, Zhifei, et al.
Published: (2024)
by: Shi, Zhifei, et al.
Published: (2024)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
by: Lu, Weiheng, et al.
Published: (2024)
by: Lu, Weiheng, et al.
Published: (2024)
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos
by: Zhao, Zecheng, et al.
Published: (2025)
by: Zhao, Zecheng, et al.
Published: (2025)
Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Energy-Based Open-Set Active Learning for Object Classification
by: Lyu, Zongyao, et al.
Published: (2026)
by: Lyu, Zongyao, et al.
Published: (2026)
Similar Items
-
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
by: Li, Zongyao, et al.
Published: (2025) -
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
by: Li, Wei, et al.
Published: (2024) -
Learning to Predict Gradients for Semi-Supervised Continual Learning
by: Luo, Yan, et al.
Published: (2022) -
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
by: Chai, Zenghao, et al.
Published: (2024) -
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
by: Guo, Yangyang, et al.
Published: (2023)