Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Chaochen, Luo, Guan, Zuo, Meiyun, Fan, Zhitao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
von: Yu, An, et al.
Veröffentlicht: (2025)
von: Yu, An, et al.
Veröffentlicht: (2025)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
von: Park, Seojeong, et al.
Veröffentlicht: (2024)
von: Park, Seojeong, et al.
Veröffentlicht: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
von: Cao, Zhuo, et al.
Veröffentlicht: (2025)
von: Cao, Zhuo, et al.
Veröffentlicht: (2025)
SLVideo: A Sign Language Video Moment Retrieval Framework
von: Martins, Gonçalo Vinagre, et al.
Veröffentlicht: (2024)
von: Martins, Gonçalo Vinagre, et al.
Veröffentlicht: (2024)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
von: Zhang, Shihang, et al.
Veröffentlicht: (2026)
von: Zhang, Shihang, et al.
Veröffentlicht: (2026)
Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video Generation
von: Davtyan, Aram, et al.
Veröffentlicht: (2023)
von: Davtyan, Aram, et al.
Veröffentlicht: (2023)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
von: Um, Sung Jin, et al.
Veröffentlicht: (2025)
von: Um, Sung Jin, et al.
Veröffentlicht: (2025)
From Sora What We Can See: A Survey of Text-to-Video Generation
von: Sun, Rui, et al.
Veröffentlicht: (2024)
von: Sun, Rui, et al.
Veröffentlicht: (2024)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
von: Kim, Ji-Hyeon, et al.
Veröffentlicht: (2026)
von: Kim, Ji-Hyeon, et al.
Veröffentlicht: (2026)
Structure-Aware Prototype Guided Trusted Multi-View Classification
von: Huang, Haojian, et al.
Veröffentlicht: (2025)
von: Huang, Haojian, et al.
Veröffentlicht: (2025)
Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2023)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
Guess the Unified Model: How Much Can We Recover from Generated Images?
von: Cekinmez, Jasin, et al.
Veröffentlicht: (2026)
von: Cekinmez, Jasin, et al.
Veröffentlicht: (2026)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
Can We Change the Stroke Size for Easier Diffusion?
von: Bai, Yunwei, et al.
Veröffentlicht: (2026)
von: Bai, Yunwei, et al.
Veröffentlicht: (2026)
Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
von: Yang, Junkai, et al.
Veröffentlicht: (2026)
von: Yang, Junkai, et al.
Veröffentlicht: (2026)
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
von: Fan, Qi, et al.
Veröffentlicht: (2024)
von: Fan, Qi, et al.
Veröffentlicht: (2024)
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
von: Wang, Ni, et al.
Veröffentlicht: (2024)
von: Wang, Ni, et al.
Veröffentlicht: (2024)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
von: Long, Xiaosheng, et al.
Veröffentlicht: (2025)
von: Long, Xiaosheng, et al.
Veröffentlicht: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)
eDOC: Explainable Decoding Out-of-domain Cell Types with Evidential Learning
von: Wu, Chaochen, et al.
Veröffentlicht: (2024)
von: Wu, Chaochen, et al.
Veröffentlicht: (2024)
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
von: Zhu, Zihao, et al.
Veröffentlicht: (2026)
von: Zhu, Zihao, et al.
Veröffentlicht: (2026)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
von: Vu, Huu-An, et al.
Veröffentlicht: (2025)
von: Vu, Huu-An, et al.
Veröffentlicht: (2025)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
von: Wu, Peng, et al.
Veröffentlicht: (2023)
von: Wu, Peng, et al.
Veröffentlicht: (2023)
Maintaining User Trust Through Multistage Uncertainty Aware Inference
von: Agrawal, Chandan, et al.
Veröffentlicht: (2023)
von: Agrawal, Chandan, et al.
Veröffentlicht: (2023)
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection
von: Yang, Jin, et al.
Veröffentlicht: (2024)
von: Yang, Jin, et al.
Veröffentlicht: (2024)
SpaceSense-Bench: A Large-Scale Multi-Modal Benchmark for Spacecraft Perception and Pose Estimation
von: Wu, Aodi, et al.
Veröffentlicht: (2026)
von: Wu, Aodi, et al.
Veröffentlicht: (2026)
CapHuman: Capture Your Moments in Parallel Universes
von: Liang, Chao, et al.
Veröffentlicht: (2024)
von: Liang, Chao, et al.
Veröffentlicht: (2024)
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video
von: Zeng, Huajian, et al.
Veröffentlicht: (2026)
von: Zeng, Huajian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
von: Yu, An, et al.
Veröffentlicht: (2025) -
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
von: Park, Seojeong, et al.
Veröffentlicht: (2024) -
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
von: Yuan, Huaying, et al.
Veröffentlicht: (2025) -
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
von: Cao, Zhuo, et al.
Veröffentlicht: (2025) -
SLVideo: A Sign Language Video Moment Retrieval Framework
von: Martins, Gonçalo Vinagre, et al.
Veröffentlicht: (2024)